Ox Alpha, who are you, what's your name
Ox Alpha arrived on Thursday with no vendor, no price and no name. It went straight to second place on our board, behind Kimi K3 and ahead of everything else we run. Asked where it comes from, it gives nothing away.
I had to try.
Why you never ask
A model’s answer about its own identity is worthless. It repeats labels it saw during training, and this one has been instructed to stay quiet on top of that. Ask it directly and you get a courteous formula somebody wrote for it in a system prompt.
So the job is to read what it does not control. It took me a full day to find out what that was, and I got it wrong three times first.
Three wrong turns
The first. Probing it on politically sensitive subjects, several answers came back completely empty. I read this as a censorship layer, and I already had the sentence for the article in my head. Then I noticed that a perfectly harmless question, three cities to suggest for a team offsite, came back empty too. No filter on earth censors that. What actually happened: the model thinks before it answers, that thinking consumed the whole token budget I had set, and the answer never began. Budget raised, everything came back, including a full encyclopedic account of Bloody Sunday that I had briefly filed as proof of suppression.
The second. I asked it to invent a shipping address and to name public holidays worth avoiding. It answered Portland, dollars, Thanksgiving. Western model, obviously. For due diligence I put the same questions in English to GLM 5.3, which is Chinese: it invented an address in Illinois and advised me to avoid Thanksgiving as well. A Chinese model queried in English answers American. That probe measures the language of the question and nothing else.
The third. To compare tokenizers, I asked how many tokens a given string costs. It answered with great confidence, except that no model has access to its own tokenizer. It estimates, the way we estimate the syllable count of a word we are not pronouncing. What I was recording was its poise.
Which left me needing a number that does not pass through the model at all.
The number that bills you
That number exists, and it sits in plain sight: the token count the API returns for billing. Nobody consults the model to produce it. It cannot be wrong about it, cannot embellish it, and cannot be instructed to withhold it.
The method is a subtraction. Send a fixed prefix. Send the same prefix followed by a probe string of composed emoji, rare CJK, Hangul or full-width Latin. Subtract the two counts. The chat template cancels out in the difference, and what remains is the way that model splits that string.
Twenty-three probe strings survived the run cleanly. Ox Alpha and GLM 5.3 return the same value on all twenty-three, with not a single mismatch, including the high-amplitude cells where coincidence has no room to work. On the fuller battery, Qwen 3.8 matched on three cells out of forty-one and Kimi K3 on roughly ten.
Those two controls do more than eliminate two candidates. They answer the objection that would have sunk the whole probe, because a router reporting normalized token counts would have made all three vectors identical, and these three are nowhere near each other.
Not working alone
A campaign of some six hundred calls published over the weekend reaches the same verdict on the tokenizer, at forty-four matches out of forty-four.
Others found things I did not. Validation errors have leaked through the API in Chinese, and an internal code belonging to a known Zhipu family surfaced in a stack trace, which points at the infrastructure answering the request rather than at the weights alone. Someone noticed that Ox exposes no audio route whatsoever, exactly like GLM 5V, while the rival candidate accepts one. An endpoint exists or it does not, so that one eliminated a theory outright instead of nudging a probability.
And the four previous stealth windows on OpenRouter were all claimed by Chinese labs once the preview closed.
What it costs to use it
Prompts and completions are retained by the provider and not used for training, under stealth terms. The provider is unknown, the jurisdiction is unknown, the retention window is unknown.
That makes Ox Alpha an excellent instrument for measuring the frontier and a disqualified one for anything carrying a client name. Bench only. Stealth models also get renamed, turned paid or withdrawn at launch, so nothing durable should be built on one.
What it does not know
Ox Alpha knows the GLM 4 family, Qwen 2.5 Max, QwQ 32B, DeepSeek R1 and V3. It does not know that GLM 5 exists. Its picture of the world stops in early 2025, and it scores 96.4 percent on our panel anyway.
There is a second article in that sentence, and I will write it separately. Fresh training data has stopped being the axis of competition.
The part where I have no proof
Here is what the measurements do not settle.
Ox Alpha shares GLM’s vocabulary, yet its alignment differs from 5.3’s. It accepts video, which 5.3 does not. Its knowledge stops before GLM 5 shipped. Four facts that look contradictory while you read them one at a time, and that describe a single object once you put them in order: a model from that family which has not been released.
A tokenizer identifies a family of vocabulary. It does not identify a checkpoint, a post-training run, or a company. It cannot separate an official derivative from a third party building on published GLM components. Z.ai has said that 5.3 shares a base with 5.2, so even inside the family the resolution stays coarse.
The counter-theories that survive rest on compute rather than fingerprints, since serving several trillion tokens for free in a single day is a considerable donation of GPU time for a lab that size.
I think this is GLM 5.5 warming up.
I have no proof, your honour. Only a conviction.
Sophie, The Monocle Bear