GLM 5.3 Flash : On prem, as premium
GLM 5.3 Flash landed on Hugging Face this Friday night. I spent the weekend running it. Six hundred thousand tokens later, I have its picture: a must-have locally, and this time not a model built for a datacenter.
It is a sharpened version of the 5.3, 320 billion parameters with 18 billion active on any given token, and its real argument is size. In full BF16 it weighs 583 GB, too much for the machine. In 8-bit with the heads kept in BF16 it drops to 311, with near-total perplexity recovery, and that bar clears. Z.ai put two builds on the Hub, the full BF16 and an FP8, and the FP8 is very likely what served the cloud preview. FP8 and my Int8 are not the same quantization, but they sit close enough that the local model and the cloud model are, to a first approximation, the same weights. That is the whole reason the comparison is worth making.
A week ago an anonymous model called OX-Alpha had shown up, scoring in the leading pack of my bench. I had it down as GLM from its fingerprint and bet on the long-promised 5.5. It was 5.3 Flash instead, reason enough to spend the weekend on it.
A Saturday morning of code before a single token
The trials were not smooth. The architecture was not yet supported in MLX, so I had to write it. That took the Saturday morning, not as a download bar filling up but as errors falling one by one. Once they were gone, the tests ran.
Decode speed settled around twenty-two tokens a second. For a 320B model on local hardware, that is honorable, and I took it as a win.
The win lasted until the first real test.
The bill comes due in reasoning
The model is verbose. More verbose than a drunk uncle ranting through an entire Xmas dinner. Its reasoning does not seem to know how to stop, looping back, restating, filling the window.
The 96,000-token ceiling I usually set was not enough. I had to lift it to 300,000 to give it room to finish its homework, and even then I ended up cutting the reasoning effort on a couple of tests just to get out before the night was over. This is the caveat that does not appear on any leaderboard: thinking, done locally, weighs the inference down. It is the price of the quality, paid in wall-clock time.
The score is the budget
I ran the same coding benchmark twice in one day, same model, same harness, changing only the reasoning effort. At reduced effort it scored 86.4. At maximum effort, 92.0. Five and a half points, moved by one variable, the token budget granted for thinking. The cost of those points was nearly three times as many tokens produced.
The tests that climbed the most were the ones whose failure at low effort was mechanical, not conceptual: a broken awk quote, an axis it had forgotten to invert, a regex that matched no label at all. Nothing that looks like missing skill. Just reasoning cut off too early. Given the room, it repairs them one by one.
This model is not mediocre. It is good when you let it think. A score measures the budget you granted as much as the engine you ran, which is worth remembering in front of any benchmark published without a word on the configuration behind it.
It fits. It just does not have the time.
There are two ways to fall short of the cloud, and they do not get fixed the same way. Kimi K3 does not fit in the machine at all. That is a wall of capacity, and nothing short of new hardware moves it. GLM 5.3 Flash fits, runs, produces. Its wall is throughput. OpenRouter decodes the same model at roughly a hundred tokens a second, a factor of four and a half over my local rate. On an engine that has to emit twenty thousand tokens of reasoning per test, that factor stops being a comfort detail and starts deciding what the model has time to finish.
On my panel the local build scores 95.2 against 97.6 for the cloud version. The gap lives almost entirely in code and planning, the trials that demand the longest reasoning, which is exactly where the missing throughput bites.
My default coding model, with an asterisk
On OpenRouter, GLM 5.3 Flash is now the model my coding harness calls most of the time. The token is cheap there, in and out, and for this level of quality that is what wins the daily decision.
On the same 32 tests, hosted, Flash and Kimi K3 finish a tenth of a point apart, and Fable 5 a point below Flash. Then the invoices arrive. Put Flash next to Fable 5 and the bill divides by ninety.
| Model | Output tokens | Total cost | Cost per test | Price factor |
|---|---|---|---|---|
| GLM 5.3 Flash | 596,312 | $0.15 | $0.0048 | ×1 |
| GLM 5.3 | 264,752 | $1.15 | $0.036 | ×7.6 |
| Kimi K3 | 601,853 | $7.77 | $0.243 | ×51 |
| Fable 5 | 253,792 | $13.21 | $0.426 | ×90 |
Flash and Kimi K3 emit almost the same volume, around six hundred thousand tokens across the panel, so the entire gap in their bills is price, not behavior. That is the asterisk on the cheap token. Flash is a chatterbox, roughly twice the output of the terser engines, and only a rock-bottom unit price absorbs it. On a model that charged more per token, that verbosity would be the whole invoice.
And I owe one honest line about what that 95.2 measures. Not the cloud model, but a build I ported to MLX by hand, squeezed to 311 GB, and judged on a panel I had to cut short for lack of machine time. That OX-Alpha and this release are the same model is a fingerprint correlation, not a signed fact. Saying so does not weaken the result. It is the condition for the result to mean anything.
So, is the local up to par? Yes, at a price paid in patience.
Sophie, The Monocle Bear