Configuration Parsing Warning:In config.json: "quantization_config.bits" must be an integer

Yamz Labs

GLM-5.3-Flash, EXL3

Model size: 320B parameters (18B active). The "50B params" figure in the Hugging Face sidebar is not a parameter count: the site counts the packed 16-bit words of the EXL3 format.

Hardware: sized to fit one 128 GB machine. The pack is in the standard EXL3 format, so it is not tied to one GPU; it has been tested only on AMD Strix Halo (Ryzen AI Max+ 395) with the Kyojin engine, and other hardware is untested.

An EXL3 quantisation of zai-org/GLM-5.3-Flash (MoE, 320B total, 18B active) for one 128 GB unified-memory machine (AMD Ryzen AI Max+ 395, Radeon 8060S, gfx1151). It includes the MTP head for speculative decoding. It runs on the Kyojin engine, built on ExLlamaV3, not on stock ExLlamaV3. Quantised tensors by turboderp; layer mix, scale-vector tuning and packaging by Yamz. Not affiliated with Z.ai or turboderp.

Highlights

  • Quality. Against the official FP8 model, KLD (how far the output distribution is from FP8; lower is closer) is 0.151 and top-1 agreement is 89.31 % on 30 held-out rows. On 126 of the 129 held-out rows top-1 agreement is 90.11 %.
  • Size. 99.73 GB for a 320B-total MoE: one 128 GB machine holds it, with MTP and long context.
  • Speed. Prefill (reading the prompt) 580 tok/s at 3.5K, 584 at 14K and 546 at 64K; decode 26 to 30 tok/s with MTP (2 drafts). One Ryzen AI Max+ 395.
  • MTP projection included. The pack ships mtp_eh_proj.st (64 MB), an unquantized copy of the MTP projection. The Kyojin server loads it automatically from the model folder (engine main from 2026-10-05, no flag). It only changes the draft, not the model's answers: 31.6 tok/s greedy with it against 26.7 without, on three prompts of 400 tokens.
  • Engine. Kyojin, the Yamz engine built on ExLlamaV3, runs it on gfx1151 with ROCm, MTP head included.
  • Feedback wanted. Run it on your Strix Halo / Ryzen AI Max machine and tell us your tok/s and hardware: open an issue at https://github.com/Yamz-Labs/kyojin/issues or start a discussion on this page. GLM-5.3-Flash on the same machine: Yamz engine vs llama.cpp KLD and top-1 agreement against the official FP8 model

How it was made

The quantised tensors come from turboderp's public EXL3 quantisations of this model (turboderp/GLM-5.3-Flash-exl3): the 2.05 bpw pack, with the expert layers 29 to 44 taken from the 3.05 bpw pack. Yamz chose that per-layer mix, then ran one short tuning stage: small per-layer scale vectors (su/sv) of layers 21-44 are tuned against the official FP8 model on 526 text rows of our own. The tuned vectors are written into the normal shards: the pack is a plain EXL3 pack with 12 shards and no override file. Size: 99.73 GB.

Run

git clone https://github.com/Yamz-Labs/kyojin && cd kyojin
./build.sh                                   # builds the ROCm extension for gfx1151
source tools/strix_halo/env.sh
python tools/glm/serve.py --model /path/to/this-pack --port 8000 -c 131072 --num-draft 2

OpenAI-style HTTP API. --num-draft 2 turns on MTP. -c 524288 loads (see Hardware). Exact flags: tools/glm/SERVE.md in the engine repo. These steps were run from a clean install on 2026-10-05.

Hardware

Machine Status
Ryzen AI Max+ 395, 128 GB LPDDR5X, Radeon 8060S (gfx1151), ROCm Tested. All numbers below.
Other AMD GPUs, NVIDIA GPUs, less than 128 GB Not tested. The pack is 99 GB: it does not fit in less memory.

Memory: at -c 524288 the server holds about 118 GiB of the 124 GiB the system shows, and leaves 1 to 5 GiB free. Treat 128 GB as the floor for this pack at that context: close every other large process, and start clients outside any memory-limited session. Lower -c reduces the KV cache; the footprint at -c 262144 was not measured.

Quality (reference = official FP8)

Rows Engine KLD vs FP8 Top-1 agreement with FP8
30 held-out rows, 30,720 tokens this repository's engine (stock path) 0.15081 89.31 %
4 long rows of 4,096 tokens, never trained on this repository's engine (stock path) 0.22737 88.74 %
126 of the 129 held-out rows (the longest were left out by a memory guard) served engine not run 90.11 %

The 30 held-out rows and the long rows are disjoint from the training rows (exact-row and 12-gram token-shingle checks against all 129 held-out rows and the long rows; 36 of our rows that shared any 12-gram were removed from training). We report no figure measured on our own training domain. KLD on all 129 rows was not run. The shipped pack (12 shards, 151,554 tensors) reproduces the 30-row result bit for bit on the stock engine: KLD 0.150812.

Quality is measured against FP8, not against other formats. No same-size GGUF quality comparison was run. Task scores of this exact pack were not run; reasoning mode was not part of any task test (non-thinking mode only).

Speed (served, MTP 2 drafts)

Every row says which machine, sampling and hook state it used. "Hook" is the optional runtime steering hook of the engine (see the presets repository). "Agent flags" are the server flags of our own agent setup: chat template with medium reasoning effort (--chat-template, --default-reasoning-effort) and --max-history 2. All speeds come from development builds of the engine; the published tree differs from them in a few kernel (including the MoE prefill kernel) and quantiser source files and was not itself benchmarked for GLM.

First launch. The engine tunes its dense GEMM kernels on the first requests and keeps the result in a cache. On a fresh install the first GLM prefills run at 200 to 240 tok/s; speed reaches the figures below within a few requests and stays there on later launches.

Hook on, agent flags on. Ryzen AI Max+ 395, 128 GB, server at -c 98304, MTP 2 drafts. Prefill, mean of 3 runs (standard error): 3.5K prompt 580 (10), 14K 584 (3), 64K 546 (1.5) tok/s. Decode, 128 new tokens, mean of 6 runs (standard error), prose / chat / code: with client temperature 0, 29.0 (0.7) / 30.3 (0.4) / 27.6 (0.6) tok/s; with the server default sampling (temperature 1.0, top-p 0.95), 26.0 (0.8) / 28.5 (0.7) / 26.4 (0.9) tok/s. No run at 128K context.

Single runs.

Setup Prefill tok/s (prompt size) Decode tok/s
-c 131072, 24K prompt 609.0 28.1 (28.12, 28.09)
-c 524288, raw engine server, steering preset loaded 611.9 (3.7K), 585.0 (14.2K) 27.1 default sampling, 28.1 greedy
-c 524288, through an agent gateway (OpenAI-style proxy in front of the server) 590.0 (3.7K), 589.8 (14.4K); 22.8K cold prompt: first token after 39.0 s, about 585 29.1 default sampling, 28.1 greedy

The proxy path decodes 1.07 times the raw figure on default sampling: run-to-run variation, not a speed-up. A 185K-token recall test found its needle, but the cold prefill of that prompt was not timed. Context settings differ between rows (-c 98304, -c 131072, -c 524288): compare rows within the same setup. Repetitions inside one setup differ by up to 20 %, so no difference below about 10 % is resolved; the hook alone was inside that noise, and a microbenchmark of the hook measured -0.1 to -0.3 % of decode.

For scale: llama.cpp (ROCm, UD-IQ1_S 1.56 bpw, -fa 1 -ub 2048, no MTP) on the same machine: pp4096 197.9, pp16384 158.3, tg128 16.74, tg at 64K 7.19 tok/s. That quant is coarser than ours, so the comparison covers engine and format together.

Optional steering presets

Optional steering presets for the engine hook are published separately in the Yamz presets repository (https://github.com/yamz-labs/yamz-presets); they are not part of these weights. A ready-made variant of this pack, with the preset bundled as a file the engine applies at load, is yamz-labs/GLM-5.3-Flash-EXL3-Yamz-Uncensored (same shards).

Sampling

Use the base-model card values: temperature 1.0, top-p 0.95, no repetition penalty. Greedy decoding is not recommended.

Limits

  • Sampling. Use the card values above; greedy decoding is not recommended, especially for long outputs.
  • Only gfx1151 with ROCm is tested.
  • The conversion tooling is not published. The engine and these weights are.
  • Reasoning mode was not part of the task tests (non-thinking mode only).

Licence and credits

MIT. Base model GLM-5.3-Flash, Copyright (c) 2026 Z.AI Co., Ltd, MIT, licence file included as LICENSE. Our additions: Copyright (c) 2026 Yamz Labs, MIT. Provided as is, without warranty. Engine: built on ExLlamaV3 by turboderp (MIT). GLM is a name of Z.ai; this release is unofficial.

Training data of the base model: see the upstream card. Our tuning stage used 526 text rows of our own (mixed domains: chat, code, agent, math, French); the held-out proof is above.

Quantised tensors by turboderp; layer mix, scale-vector tuning and packaging by Yamz. Not affiliated with Z.ai or turboderp.

Downloads last month
456
Safetensors
Model size
50B params
Tensor type
BF16
路
F16
路
I16
路
F32
路
Inference Providers NEW
This model isn't deployed by any Inference Provider. 馃檵 Ask for provider support

Model tree for yamz-labs/GLM-5.3-Flash-EXL3-Yamz

Quantized
(155)
this model