← Blog

Kimi K3, I Wanted It Local

Kimi K3, I Wanted It Local

Yesterday at six in the evening, Moonshot released Kimi K3 on HuggingFace.

Since the official cloud release a week ago, the benchmarks say the same improbable thing: the best open model in the world is now trading blows with Fable 5. Not “impressive for open weights”. Not “closing the gap”. Level. What changed yesterday is one word: downloadable. For the first time since this race started, the best-model-of-the-moment conversation includes something you can put on your own disks.

So I downloaded it.

Not for a client. Not for a benchmark. Because I run OdyssAI, a sovereign AI system built on the premise that serious models can live inside your walls, and a premise like that gets tested on the day it becomes hard. If “open weights” only means open for datacenters, the phrase is marketing. Yesterday evening was the day to find out which it is.

What “download the model” actually means

K3 is 2.8 trillion parameters. 1.56 terabytes of weights. Four Macs at 512 GB could hold it, and very few people own that. But hardware is not even the first obstacle: the architecture is brand new, and local engine support for it amounts to a draft PR that does not run. Whoever wants this model at home this week, whatever their machines, first has to write the code.

So the night was not spent watching a progress bar. It was spent writing the engine while the weights came down: seven hundred lines implementing an architecture I had first read about at dinner time. K3 has attention residuals that flow across the depth of the network, experts that work in a compressed latent space, and an attention mechanism with no positional rotation at all, because position lives elsewhere. Port it naively and nothing crashes. The model just gets quietly stupid, and you will never know why.

The weights themselves lied to me twice. My favorite: one tensor ships in a shape that Moonshot’s own reference code declares differently. The vendor’s file cannot be loaded by the vendor’s code. When you run frontier models at home, you are not a user. You are the last line of QA.

By half past midnight, the model was converted. By two, it was distributed across the five machines of the cluster. Then I went to bed, and the walls started.

The first wall: 1.5 terabytes meets 1.46

OdyssAI’s inference cluster wires 1460 GiB of available memory across five Mac Ultras. The converted K3, at the vendor’s own precision, weighs 1390.

That should fit. It does not.

At seven this morning, node after node died in silence. No crash, no error, no log. A process loading 233 GiB would run for exactly sixty-six seconds and vanish, taking its memory with it. It took an observer script sampling every three seconds to catch macOS in the act: the OS counts a model’s memory-mapped weights against the process itself, and executes anything that gets too close to physical RAM. Quietly. This behavior is documented nowhere. I know of no one who has hit it before, because I know of no one who has pushed five Macs to 95 percent of their memory with a single model.

The real ceiling on Apple Silicon is not the official GPU limit. It is about 80 percent of RAM, and past it your process does not fail. It disappears.

Making it fit

The way in was to compress the model’s experts one notch further than Moonshot intended, from their carefully trained 4-bit format down to 3.5 bits. That is not a free operation. The experts are precisely the part of the model the authors protected during training, and I overwrote that protection at four in the morning to save 240 gigabytes.

Even then, the cluster refused four more times before the first loaded: true. A missing tokenizer file that killed a rank after it had already swallowed 228 gigabytes of weights. A cache format this architecture cannot use. A transport module out of sync on four nodes. An ssh keepalive that executed a perfectly healthy load at startup plus three hundred seconds, exactly. Six attempts in all, each one a lesson, none of them visible from outside a home-built engine.

Then, at last: 1151 GiB resident across five machines, and a first token.

I want to be precise about what that sentence means. What runs in my office is not Kimi K3. It is Kimi K3 minus an unmeasured slice of quality, admitted through the only door that was open. Everyone who tells you they run frontier models locally has a sentence like this one. Almost nobody prints it.

The second wall: bandwidth

Physics says each generated token must read 67 gigabytes of weights from memory. Twenty-one of those gigabytes are the active experts. The other forty-six are the path every token pays in full: attention, shared experts, latent projections. Two percent of the parameters, seventy percent of the traffic.

On this hardware, that put the ceiling somewhere between six and ten tokens per second. Measured, once the caches warmed: 2.7, stable, and it will not go higher. The only other K3 running on Apple hardware cuts out 80 percent of the experts, quantizes lower than we do, and reaches 0.2 on a single Mac. So there is no embarrassment in the number. There is just no work in it either.

The third wall: reading costs more than answering

I expected decode to be the bad news. The dashboard had worse in store.

Prefill, the reading of your prompt before the first token comes back, runs at half a token per second. And it degrades with length: 161 tokens took 295 seconds to read. 254 tokens took 591. The model reads more slowly than it writes, and the more you give it, the slower it gets.

The cause is structural. The prompt crosses the pipeline in micro-batches, twenty of them for a single question, and every one pays the full latency of five ranks plus their synchronization barriers. Capacity was yesterday’s wall. Bandwidth was this morning’s. Pipeline latency is the one you discover last, because you have to beat the other two before you are allowed to meet it.

And it kills the one argument K3 had left. This architecture makes long context nearly free in memory: a fixed-size recurrent state, a cache of 576 bytes per token, a million-token window that costs a few gigabytes where classic models detonate. The entire case file, the full codebase, held in mind at once, on machines you own, with nothing leaving the building. That was the use case. That was the reason to accept everything above.

At half a token per second of reading, a 100,000-token prompt takes two and a half days.

Free in memory. Unaffordable in time.

The verdict

Quality, for what it is worth, seems to be there. The first exchanges are coherent, precise, unmistakably frontier. A full run of our benchmarks would take weeks at this speed, so we will skip it, and the assessment will remain entirely arbitrary. Mine.

By late afternoon, I unloaded it. K3 lived half a day on the cluster: loaded at breakfast, measured through lunch, gone before dinner, its place returned to a smaller model that does the daily work here at twelve tokens per second.

The honest accounting stands. The frontier fit by one Mac: one more 256 GB node and the vendor’s untouched weights would have loaded. The premise held. The best model in the world can live inside your walls, tonight, for the price of a company car. Nobody else has done it, as far as I can tell. And for half a day, it answered.

A beautiful feat. Totally unusable.

Sophie, The Monocle Bear