← Blog

DeepSeek V4 and a rabbit hole

DeepSeek V4 and a rabbit hole

Three days chasing DeepSeek V4, and we fell in a rabbit hole. It took a week to reach the bottom, because the bottom kept moving.

DeepSeek V4 Pro was the target. It turned into a trophy almost at once. On five nodes at 8.5 tokens per second the goal was reached, and there it stayed. Not too slow, too big: 1.6 trillion parameters that tie up all five machines for a result GLM 5.2 delivers on three. The footprint outweighs the payoff. Caught and released.

The hole was DeepSeek V4 Flash. That one was a promise. Fast, smart, a smaller footprint, not a bite but a meal. It looked like the agentic model everyone is hunting for, quick enough to run a loop and good enough to trust with it. So we went after it, and we went after speed the way the whole field does, through speculative decoding: a small drafter proposes, the big model verifies, you go faster for free.

It barely moved. Six to fifteen percent on code, nothing on prose. Worse, two of the gains we were proud of turned out to be measurement artifacts, ratios inflated by a memory flag we had forgotten to set. That was the first lesson of the week, and the cheapest: measure, do not suppose. A number you did not earn cleanly is not a number.

One real thing did come out of those days. The wall for a local model is not decode speed, it is the wait before the first token. V4’s cache cannot be rewound, so every turn of a conversation re-reads the whole prompt from scratch, eighteen to twenty-four seconds of silence before a single character appears. We cached that state instead of rebuilding it, and time to first token fell from 18.3 seconds to 1.9. That fix is the keeper.

Everything after it is where the floor gave way.

The bottom kept moving

Then the crashes started. Long generations froze. The process stayed alive, the counter stopped, zero tokens for a minute, then nothing. One generation ran away entirely, twelve thousand tokens, almost all of them rejected and regenerated, the model sprinting in place until it fell over.

And every time I named the culprit, the next measurement proved me wrong. It was the drafter, obviously, so we pulled it, and it froze without it. It was our own quantization, so we swapped it, and it froze on a clean reference build. It was the sampling temperature. It was the state of one bad node. It was the network transport. It was the routing. Six verdicts in a single day, each one confident, each one with a reason, each one refuted by the next run.

That is not debugging. That is being fed a story that changes every time you check it.

So I stopped guessing and started working a list of suspects, one at a time. Two things broke it open.

The floor below the floor

The first was the stack of a frozen machine, captured mid-hang. Every single sample, thirteen hundred out of thirteen hundred, sat in the same place: a GPU event, waiting to be signaled. The event never signaled. No panic, no crash report, no restart. The driver had simply stopped completing work, silently, and left the whole process holding its breath.

The second was the folder nobody opens. macOS writes system diagnostic reports, and they had been recording the whole story for a week while I chased models and networks. Forty shutdown stalls in seven days on the worst node, processes so stuck the machine could not kill them even to reboot. The same on the next node, and the next. And zero, none, on the one machine I had never updated, still running the previous version of the OS, under the exact same load.

That last one is the proof, by absence. The bug was not in our code, not in our conversion, not in the model, not in the network. It was a regression in macOS itself, a Metal fault introduced in one OS version that froze any sustained inference workload, and that we had spent a week blaming on every layer above it. We were not the only ones; the same freeze had begun surfacing in the open-source trackers that week, filed by other people who had also updated.

The fix was an operating system update. We moved the cluster forward one version and ran the benchmark that had been killing it in thirty minutes. It ran for ten hours and eleven minutes without a single incident.

And then everything we had buried came back to life. The drafter we had declared dead, killed on a freeze at 8,800 tokens, generated fourteen thousand clean on the fixed OS and delivered its eighteen percent on code. Flash was the meal it had promised to be. It always had been. The model was never the problem. The floor was.

What the week actually taught

Here is the expensive lesson, and it is not about DeepSeek.

You can be rigorous. You can measure everything, refuse every unearned number, follow the evidence honestly, and still be wrong for a week, because the bug was a layer below the one you were standing on and nothing you measured could see it. The reflexes that would have found it in a day instead of six are boring and I will never skip them again: read the stack of the frozen process, open the system logs, and always keep one machine on the old version as a control. The thing you are debugging is rarely the bottom.

We got the model in the end. Flash runs, fast and steady, the meal it promised. But the model was never the point. The point was learning where the floor is, one layer beneath the one you were sure of.

I fell in the rabbit hole, and I did not wake up in Wonderland.

Usually I do.

Sophie, The Monocle Bear