← Blog

DeepSeek V4.1 Flash: a never-ending story

DeepSeek V4.1 Flash: a never-ending story

At its maximum reasoning setting, DeepSeek V4.1 Flash produced 1.1 million characters of internal monologue overnight and never emitted a single line of the code it was asked for. The same test, same machine, reasoning switched off: 92 out of 100, in five thousand tokens, in under two minutes.

That is the most extreme version of the finding. The moderate version is worse news, because it shows up at every setting in between.

DeepSeek shipped V4.1 Flash a few weeks after GLM 5.3 Flash and Qwen 3.8 Flash Next. Three labs, three small-and-fast models, one shared promise: near-frontier behaviour at a fraction of the compute. This one clears the whole benchmark panel on a five node Apple Silicon cluster in 1 hour 49, converted to MXFP4, at 88.7.

That 88.7 is the score with reasoning off. It is the best score this model produced, and it is also the cheapest.

Getting it to answer at all

The conversion went through without drama: MXFP4 weights, bf16 attention heads, 42 tests generated in 1 hour 49 at 21.3 tokens per second. The serving layer was the problem.

V4.1 emits its closing </think> as an ordinary BPE token and never emits an opening one. Any server that detects reasoning by watching for a <think> tag sees nothing at all, and pours the entire chain of thought into the content field alongside the deliverable. The second defect was blunter: reasoning_effort is documented as an integer from 1 to 100, the request layer typed it as a string, an integer returned a 422, and the string "100" was accepted and silently ignored.

The consequence is worth stating plainly, because it is not specific to my stack. The run I collected with the thinking flag enabled was byte identical to the run without it. The flag did nothing. Anyone who benchmarked this model with reasoning on, on a local server in the first forty eight hours after release, has a reasonable chance of having measured the same configuration twice and written it up as two.

Both defects are fixed in serving 1.49.2. No score labelled “with thinking” existed for this model in local before 12 September.

What four bits actually cost

Comparing a quantised local build against a cloud endpoint is usually an argument about compression. Here it turned out to be an argument about token budget.

Eight common code tests, same prompts, same judge:

runtokens per testscore
V4.1 local MXFP4, no thinking8,29185.3
V4 cloud, April build9,72688.3
V4.1 cloud, default effort24,36288.3
V4-0731 cloud47,39093.7
V4-0731-max cloud53,29093.3

Read the left column first. The table is a budget curve, not a version curve: 8k tokens buys 85, 10k buys 88, 24k buys 88, 47k buys 94. The V4.1 cloud endpoint spends three times the reasoning of the April V4 build and lands on exactly the same number. The “V4.1 regresses against V4” reading that suggests itself at first glance is an artefact of how much each endpoint was allowed to spend.

Which leaves three points between the local four bit build and a cloud run at comparable budget. Honest caveat: the judge noise measured on this campaign is 4.33 points per test on identical input, so three points across eight tests sits inside the margin. The defensible claim is not that MXFP4 costs three points. It is that MXFP4 costs nothing large enough to measure this way, and that the failure mode of bad quantisation is absent. No loops, no truncation, no degenerate repetition, every test terminated. That is the difference between a compression that works and one that does not, and it is visible without a judge.

The alternative is to not compress. A 6.7 bit build of this model weighs 463 GB and fits a 512 GiB machine only at a thousand tokens of context, which is another way of saying it does not run.

The agentic number is the one that holds

The panel that survives scrutiny is the executed one: 90 tasks, three repetitions each, temperature 0, mechanical scoring against final machine state. No judge, no prose, no opinion.

Local lands at 98.6 excluding one eliminated family. Cloud lands at 97.8 on the same suite, with a lower per-repetition success rate and 15 to 30 percent more tokens spent.

I am not going to claim that 0.8 points is a local win over the cloud. It is parity, which is the interesting result. On the part of the benchmark closest to what DeepSeek actually optimised this model for, a four bit conversion running on consumer silicon in a cupboard loses nothing to the API, and gets there cheaper. That claim did not need a judge to produce and does not need one to defend.

One caveat on reading DeepSeek’s own published agentic numbers against anything else: those are loop results at temperature 1, and comparing them to a single-shot code score is a category error in both directions.

The reasoning regression

Once the serving was fixed, reasoning became measurable. It is the reason this article exists.

On the agent panel, effort 25 took the score from 96.2 to 92.2 and multiplied tokens by 8.7. On code, effort 25 moved 87.0 to 86.4 for 2.7 times the tokens. Taken as aggregates, neither of those deltas clears the judge noise, and I will not present them as proof of anything.

The individual cases do clear it, because they are structural rather than statistical.

The sharpest is a tool call test. Without reasoning, the model emits a correct JSON tool call. With 138 tokens of reasoning spent on a response that needed 35, it emits the same call with the name field missing. Twenty four points, not from a worse answer, from an answer that stopped being parseable. Next to it, a clarification test where the model asks about three axes at once instead of one, and a code test where it adds a return None nobody requested.

Then the one test I ran across the full range, a refactoring task whose entire point is that the output must not change:

effortscoretokens
no thinking92.05,038
2587.019,897
5083.023,921
cloud, default88.020,631
V4-0731 cloud84.0around 47,000
100, localno outputaround 300,000

Monotonically decreasing, on local and cloud alike. At effort 50 it introduces a bug that does not exist at 25: it reaches for statistics.mean() on a list of integers, which returns an integer, so a field that must serialise as 40.0 serialises as 40. A type regression, on a test whose only requirement is byte identical behaviour.

At effort 100, the number in the last row is not a score. The model reasoned all night and stopped there.

The pattern is consistent enough to name. Reasoning helps this model when the task is open and the answer has to be constructed from nothing, which is where the two gains sit. It hurts everywhere the instruction is to stay inside a frame: refactor without changing, answer short, emit exact JSON. Every additional reasoning token is one more opportunity to improve something that was not supposed to be touched.

So the shipped configuration is reasoning off. 88.7 is the number, and it is the number from the setup that costs the least.

Twelve tests failed on a missing newline

The first pass of the agentic benchmark gave local 72.0 against cloud 80.5. That result was wrong, and it was wrong in my harness.

Nine tasks scored zero because the state comparison used strict equality and the file the model wrote did not end with a newline. Three more scored zero because the number extractor took the first integer in the final message, and the model had written that zero orders remained pending and six audit rows had been created. It took the zero. In all twelve cases the machine state was exactly correct.

Ten points of gap, manufactured by a string comparison and a regex, on a benchmark whose whole selling point is that it scores mechanically instead of asking a model for an opinion.

I want to be precise about how those bugs were found, because the mechanism is not flattering. I went looking because the local number offended me. If the local build had come out ahead by ten points instead of behind, the same two defects would have been sitting there, unexamined, and this article would have had a much better headline. A harness is only as neutral as the moment you decide to audit it, and that moment is almost always the one where the result contradicts what you wanted.

All four defects found in this campaign, two in the harness and two in the serving, push in the same direction. Every number here is a floor.

What this leaves

V4.1 Flash is the most economical model on the panel by a distance: 139,721 tokens for the full run against 433,000 for GLM 5.3 and 264,000 for Qwen 3.8. It converts to MXFP4 cleanly, it matches its own API on executed agentic work, and it answers fast enough to sit inside an interactive loop without anyone waiting on it.

It also ships with a reasoning dial that, on every constrained task I measured, does nothing but spend tokens to make the output worse, and at its maximum setting spends all night producing nothing at all.

DeepSeek built a model that wins by not thinking, and then shipped the thinking as the feature.

Sophie, The Monocle Bear