GLM 5.3: give me a reason
GLM 5.3 full, served locally across four Mac Studios, scores 92.8 on our panel. The same version through the API scores 93.8. One point apart.
That is not the interesting result. The interesting result is that the local run, one notch below maximum effort, generated more reasoning than the cloud run that was supposed to be running flat out. And that we have no way to verify what the remote server actually did.
The setup
753 billion parameters, 737 GB of weights at 8 bits with a BF16 head, distributed over MLX across four Mac Studio M3 Ultras, one with 512 GB and three with 256, for 1.28 TB of unified memory. All of it served by our own inference engine.
Against that, the same model family through the API. 32 tests across four benchmarks, aggregating 18 judged axes that will get their own article. What matters here: the protocol is identical on both sides, and the judge never sees which engine produced what.
| engine | effort | A | C | P | T | total |
|---|---|---|---|---|---|---|
| GLM 5.3 cloud | unspecified | 97.6 | 88.4 | 98.0 | 93.8 | 93.8 |
| GLM 5.3 local | high | 96.2 | 90.9 | 96.0 | 91.2 | 92.8 |
| GLM 5.3 Flash cloud | unspecified | 99.0 | 97.5 | 100.0 | 96.2 | 97.6 |
| GLM 5.3 Flash local | mixed | 97.6 | 92.0 | 96.5 | 95.6 | 95.2 |
One point does not support a conclusion
The temptation with a table like that is to write that the cloud keeps a slight edge. It does not, and the proof is internal to the campaign.
On re-judging the same outputs, individual tests moved by 4 to 16 points. One run went from 72 to 98, another from 73 to 89. The judge is a model, it has its own variance, and that variance is an order of magnitude larger than the gap anyone would want to interpret.
One point across an aggregate of 32 tests is noise. The honest word is equivalent. On benchmark C the local engine actually comes out ahead, 90.9 against 88.4, and there is no more reason to read into that direction than into the other.
Reasoning effort is not an observable variable
This is the real finding of the campaign, and it is a negative one.
The local run was set to high, one notch below maximum. The cloud run was sent without an effort parameter. The model card published by Z.ai states that the default is max, and that is the entirety of what we have: nothing in the API response confirms the setting actually applied, neither from the model provider nor from the router that served the request.
What is measured, by contrast, is unambiguous. The cloud run generated 264,752 completion tokens, the local run 307,717. Sixteen percent more reasoning on the machine we control, at the nominally lower setting.
| engine | completion tokens | wall time |
|---|---|---|
| GLM 5.3 Flash cloud | 596,312 | 4h05 |
| GLM 5.3 Flash local | 410,689 | 5h00 |
| GLM 5.3 cloud | 264,752 | 1h15 |
| GLM 5.3 local | 307,717 | -8h01 |
Two readings are available, and nothing in the data settles between them. Either the cloud was running at its maximum, in which case the two servers do not mean the same thing by the same word. Or the documented default does not apply to the path the request travelled, in which case we compared our high against an unknown setting.
Either way the practical conclusion is identical. The reasoning effort of a cloud run is not data, it is an assumption. It cannot be verified after the fact, and nothing guarantees it will be the same next month.
More reasoning does not mean a better score
The next reflex would be to treat token volume as the real variable and rank engines by what they spend. We checked, benchmark by benchmark. It does not hold.
Across four benchmarks and two models, the correlation between generated volume and score holds on half of them. On the other half the local engine burns more tokens and scores less. The clearest case: on benchmark T, Flash local consumes 185,766 tokens against 116,289 for the cloud, sixty percent more, and still loses.
Reasoning volume does not predict the score. It predicts the invoice.
The only comparison that stays sound is the one made all else equal: same model, same machine, same harness, one slider moved. Taking effort from low to max on benchmark C of Flash, on our own engine, yields 5.6 points, from 86.4 to 92.0, for 2.8 times the tokens.
Five point six points for a setting. One point between local and cloud. The reasoning slider weighs more than the infrastructure choice, and it is precisely the parameter public tables leave out.
What you measure when you measure an API
It is worth being clear about what a cloud run observes. You are not testing a model, you are testing a service.
Batching, routing, the exact version served that day, the quantization the provider applied, the default settings: none of it is observable, and all of it can change without notice and without a change of name. Explaining a one point gap by the provider’s internal cooking would be speculation.
The measurable fact stops at one point, below judge noise. The rest belongs to a box we do not open.
That is also why measuring locally has value beyond the score. On our machines the quantization is known, the version is frozen, the harness is ours, the effort setting is the one we wrote, and a run repeated in six months returns the same result. A cloud benchmark is reproducible only by accident.
Two currencies
Reasoning costs on both sides. It simply is not billed in the same unit.
In the cloud it is charged as output tokens at full rate, even though the overwhelming majority of those tokens will never be read by anyone. On this panel the full version costs 1.152 dollars for 32 tests against 0.152 for Flash. Seven and a half times more expensive, for 3.8 points less.
Locally it is paid in machine hours. Five hours for Flash, eight hours and one minute for the full version, in series, on hardware already owned.
The full model is in fact considerably less verbose: 307,717 tokens against 410,689 for Flash on the same series. It reasons shorter. But it decodes twice as slowly, 10.7 tokens per second against 22.7, and the saved verbosity does not make up the throughput deficit.
The full version has no case
The family ranking is unambiguous, and it is the same on both sides of the boundary.
Flash beats the full version locally, 95.2 against 92.8. It beats it in the cloud, 97.6 against 93.8. It decodes twice as fast and costs seven and a half times less. On this panel, the 753 billion parameter model wins on no dimension at all.
Widened to the rest of the panel, the price spread turns absurd:
| model | score | dollars per test |
|---|---|---|
| Kimi K3 | 97.7 | 0.243 |
| GLM 5.3 Flash | 97.6 | 0.0048 |
| Fable 5 | 96.7 | 0.426 |
| GLM 5.3 | 93.8 | -0.036 |
The first leads the second by two tenths of a point, which is noise, for fifty one times the price.
What we did not measure
Two reservations, published as they stand.
8-bit quantization probably degrades the local model slightly. The defect profile is consistent with it: these are not derailed chains of reasoning but factual recalls that slip, a Swift API that does not exist, an invented matplotlib method, a logging function name just off target. Binary errors, the kind quantization produces.
Verifying that properly means comparing the two quantizations at the exact points where the local engine went wrong, which is targeted forward work rather than an aggregate measure. Average perplexity over a corpus is dominated by easy tokens, and these defects live in the tail, where the mean erases them. On Flash, whose full precision fits in memory, it is doable. On the full model it is not.
Effort was also not homogeneous across the campaign. The local run of the full model was set to high, the Flash run to a mixed configuration. On the cloud side, the measured gaps assume the server ran at its documented default, and that assumption is not verifiable from outside. A full Flash run at max is under way to lift at least the local reservation.
What to take from it
The local versus cloud argument is emptying out. On an open weight model, with the same version of the weights, the quality gap has dropped below what a judge can distinguish. What remains is machine hours on one side and an invoice on the other, plus a difference the score does not capture: knowing exactly what produced the result.
The argument that should replace it is about configuration. A reasoning model does not have a score, it has a curve of score against effort, and publishing one point of that curve without saying where it sits amounts to publishing a number without its unit.
A benchmark without its reasoning configuration is not a result, it is an anecdote. And on an engine you do not control, that configuration is not published. It is assumed.
Sophie, The Monocle Bear