Agentic AI @ 2 FPS
Audience: Engineers watching open models converge with hardware they can own.
Reading time: ~20 minutes.
In 1998 I was a young software engineer at Microsoft, writing DirectX code on a monstrously expensive 3D accelerator that rendered simple scenes at 2 frames per second. It felt like bullshit, and I was the one writing it. Then it shipped. Consumer cards ran the same code at 30 FPS for $125, and the accidental Moore’s Law that gaming volume financed kept compounding from there: fixed-function pipelines became programmable shaders, researchers hijacked the shaders for general math, Nvidia formalized the hijack as CUDA, and CUDA caught crypto, then deep learning, then the training runs behind the LLM I now type wishes into. The entire AI era sits on silicon that exists because tens of millions of people wanted 3D games on Windows to look better.
Today, AI enthusiasts are running a 2-bit quant of a 753-billion-parameter model at 3 to 9 tokens per second, and they’re having the 1998 feeling. That’s this era’s 2 FPS. The feeling of absurdity is what building for hardware that does not exist yet feels like from the inside.
A while back, in LLM Inferencing Costs are Going to $0, I argued that the marginal price of everyday intelligence was headed to zero and that productized local-model solutions would show up soon. That post was about price; this one is about ownership, because sometime this summer we crossed a border that price doesn’t describe: frontier-class AI stopped being only a thing you rent and became a thing you can own. Not comfortably, not cheaply, and at 2 FPS. But the border is behind us, and everything interesting follows from that.
The Border Is Behind Us
My capability threshold is frozen on purpose: “Opus 4.6-class,” meaning the model Anthropic shipped in February 2026. I picked it because 4.6 is the oldest model that still feels sufficient to me for serious agentic coding work, and I keep meeting engineers who say the same. The frontier has moved well past it since (Fable 5 arrived in June) and I use the new stuff daily. That’s a fact about the frontier; it doesn’t move the threshold. When Fable 5 shipped I ran a controlled model swap on an unchanged harness to check whether it should: First Fable (which I felt was a step up, not a step function). And if your instinct is “so what, the good stuff is an API call away,” that’s exactly the point: the frontier is the thing you rent. The threshold marks the intelligence you can own, and everything that follows is about the difference.
In June, Zhipu’s GLM-5.2 reached that threshold in the open. The facts, from the model card: 753 billion parameters, about 40 billion active per token, plain MIT license, a million tokens of context. On Zhipu’s own benchmark table it clears Opus 4.6 on the hard coding evals and closes in on 4.7. Vendor numbers, so discount accordingly. The independent datapoint is better: the UK’s AI Security Institute measured GLM-5.2 performing about level with Opus 4.6 on their cyber evaluations, and open-weight models as a class trailing the closed frontier by four to seven months, down from six to ten a year ago. A government security lab now measures the gap between the frontier and the commons in months.
Two things about GLM-5.2 matter more than the benchmarks. It ran inside Claude Code, Cline, Roo, and Goose on day one, because the ecosystem standardized on frontier-shaped APIs; the models are becoming drop-in parts. And the 4-bit build, roughly 400GB, fits a 512GB Mac Studio, where a community benchmark measured 17.7 tokens per second. The 2-bit build squeezes into a 256GB machine at 3 to 9 tok/s with visible quality loss. That is the 2 FPS experience, and it’s real, replicated, and MIT-licensed.
Then the cadence got silly. On July 16, Moonshot shipped Kimi K3: 2.8 trillion parameters, 16 of 896 experts active per token, trained quantization-aware at 4-bit so its native form is already 1.4TB. Moonshot claims the thing open models have always failed at, long-horizon agentic work: on SWE-Marathon, a multi-day agentic benchmark, its table shows K3 edging out Opus 4.8 and comfortably clear of Fable 5. Hold that one loosely. The claim comes from Moonshot’s own harness while the rivals ran on others’, and the same frontier model Moonshot credits with a near-tie scores far lower on the benchmark’s official leaderboard. The mismatch tells you as much about the state of benchmarking as about the model. On general capability the picture is calmer and better sourced: independent indices put K3 just behind Fable 5 and GPT-5.6 Sol, about where Moonshot’s own launch copy puts it (“still trails”), and Nathan Lambert calls it “the strongest open model ever released.” The weights landed on the promised day, July 27, at full native precision, all 1.4TB of MXFP4.
Three days after K3, Alibaba previewed Qwen3.8-Max at 2.4 trillion parameters, claiming second place behind only Fable 5, open weights promised. DeepSeek’s V4 line shipped in April under MIT at 1.6 trillion. Tencent’s latest is Apache-licensed. Three Chinese models over two trillion parameters, all open or promised open, inside five weeks.
Whose Brain Is It?
It’s worth being precise about who “they” are, because the shorthand I keep hearing, that the Chinese government gives away frontier models, is wrong in an interesting way.
The labs are private companies having extremely capitalist years. Zhipu listed in Hong Kong in January and the stock is up roughly 2,400 percent since. Moonshot raised $2 billion at a $20 billion valuation in May and is reportedly negotiating its next round at up to $50 billion ahead of an IPO, on annualized revenue that tripled to $300 million between March and June. These are venture-backed firms making the classic second-place move: open the weights, commoditize the leader’s margin, recruit the commons as free distribution and free QA. Nobody open-weights a decisive lead.
What the state supplies is everything around the labs. The State Council’s “AI+” plan makes open-source AI an explicit national directive, with adoption targets written down like grain quotas: 70 percent of key sectors by 2027, 90 by 2030. Seventeen-plus city governments hand out compute vouchers worth up to $280,000. Beijing subsidizes API access. Xi endorsed open-source diffusion by name at July’s World AI Conference. The government doesn’t write the models. It pays for the gym, the coaches, and the plane tickets, and it has made giving the results away a matter of national strategy. Meanwhile there is no US-origin frontier-parity open-weight model. None. The export-control irony writes itself: Washington built a dam and got a floodgate.
The uncomfortable part of depending on this pipeline is that it has an off switch, and the off switch is in Beijing. Open-model progress is a strategy, and strategies get discontinued: genuine parity would end the cadence from one side, competitive collapse from the other, and an IPO-minded CFO could end it from inside. The CFO is already warming up. For you, me, and any company that isn’t selling model hosting, the Kimi K3 License is functionally MIT: run it, fine-tune it, sell what you build on it. But a model-hosting business clearing $20 million in revenue now needs a separate agreement with Moonshot before commercial use. The commons still gets the brain free, and the hosting industry just got a bill. Then there’s the state’s own hand. China’s commerce ministry is reportedly consulting on export controls for model weights themselves, up to and including a ban on publicly releasing the most capable ones. The most concrete motion toward weight export controls anywhere in the world right now is Beijing considering whether to stop giving the brains away. K3’s weights went out while that consultation was open. The next release runs the test again.
The Machines Went Backward
Here’s where the story stops rhyming with 1998 and starts running in reverse.
In March 2025 Apple shipped the 512GB M3 Ultra Mac Studio, the machine that made all of this locally hostable in the first place. In March 2026 the 512GB option quietly disappeared from the store; no announcement, just a missing configurator row. On May 5 the 256GB option followed. The biggest Mac you can order new today has 96GB of unified memory. The window in which you could buy a new Mac that fits GLM-5.2 lasted one year, and it closed three months before GLM-5.2 shipped.
Look at the orange line. It drops.
The used market did what markets do. Secondhand 512GB M3 Ultras now list around $24,000 to $28,000 against a $9,499 original price. Asking prices, not sold prices, but the direction is unambiguous. This is a consumer computer appreciating like a crypto-era GPU, because it turned out to be the last of its kind.
The reason is the memory supercycle. Conventional DRAM contract prices rose 18 to 23 percent in Q4 2025, then a record 90 to 95 percent in Q1 2026; PC DRAM alone crossed 100 percent in a single quarter. LPDDR5X, the class Apple builds unified memory from, rose roughly another 80 percent in Q2 (the 12GB-module print went from $77 to $146). The Q3 forecast, 13 to 18 percent, is what passes for cooling. Gartner’s full-year view: memory and SSD prices up 130 percent through 2026, PC prices up 17, PC shipments down 10.4 percent, the steepest contraction in a decade.
The underlying mechanism is simply market-based allocation. A gigabyte of HBM consumes three to four gigabytes of commodity DRAM wafer capacity, AI will eat about a fifth of global DRAM wafer output this year, and the fabs, led by Samsung’s HBM4 line conversions, chose the datacenter. The consequences have names. Micron shut down Crucial, its consumer brand, to concentrate on AI customers. Nvidia raised the DGX Spark from $3,999 to $4,699 mid-cycle, citing memory costs; that works out to about $37 per gigabyte. Even the ~$97K DGX Station ships with seven of its eight HBM stacks enabled, salvaged silicon at six figures, because right now not even Nvidia can get enough memory for Nvidia.
Two numbers decide whether a model runs on your desk: whether the weights fit, and how fast you can read them. Capacity and bandwidth. Generation speed is roughly bandwidth divided by the bytes you touch per token, times an efficiency factor of one-half to three-quarters. Mixture-of-experts is what makes desk-scale frontier models possible at all: GLM-5.2 reads only its ~40B active parameters each token, about 20GB at 4-bit, so the M3 Ultra’s 819GB/s of memory bandwidth yields high-teens tok/s instead of low single digits.
Here’s the inversion. In 1999, consumer volume financed the silicon, and the hardware rose to meet the software. In 2026, the datacenter outbids you for the wafer, so the software has to shrink to meet the machines, and remarkably the models are volunteering: K3 was trained quantization-aware precisely so that its ship-quality form is the small one. As for when the squeeze ends, Morgan Stanley sees prices peaking around Q4 2026 and declining from late 2027. Intel’s CEO relays what suppliers told him: “no relief until 2028.” SK Hynix’s CEO, on the day his company listed on Nasdaq, said 2027 will be the worst supply year in the industry’s history and demand will outrun supply beyond 2030. He would say that, and he also might be right.
The $125 Card Is Coming
And it’s a box of some sort: local hardware priced for a developer or a home-office server, not for a rack. I want to be careful with that sentence, because the conclusion of this essay leans on it, and the home server is a graveyard of a category. Microsoft shipped Windows Home Server and buried it. Apple runs your smart home through the Apple TV on your shelf, a server so modest most owners don’t know they operate one. The Plex box never left the hobbyist’s closet. Every serious run at the category failed the same way: files, photos, backups, and movies ride the wire just fine, so the cloud did the job without the fan noise or the admin, and the box became a chore you could cancel.
That history looks like a refutation until you notice what it selects for. The cloud won every workload that could leave the building. Consumers kept buying serious local compute anyway, wherever the workload couldn’t leave: the game console and the gaming PC are mass-market local supercomputers, bought because latency doesn’t ride the wire. And latency isn’t the only thing that pins a workload in place. Law firms are already buying maxed-out Mac Studios to run retrieval over privileged case files, because “the client documents are on someone else’s computer” is a sentence no managing partner wants to say.
The bet underneath this essay is that AI has workloads like that, several of them, and the back half of this piece is spent naming them. It’s also possible the winning box never says server on it. Local machines win mass markets by hiding inside appliances bought for something else, the router, the console, the cable box, and the first AI boxes are already selling as developer workstations. If a mass version arrives, it will be bought for what it does, and the server inside will be as invisible as the Linux in your router.
Whatever that box ends up being called, the industry is already building it. Against the memory squeeze, the roadmaps are almost comically aggressive; most of what follows is announced or reported rather than shipped, and I’ve flagged the rumor-grade parts.
Apple, per Gurman’s July reporting (single-source, rumor-grade): a Mac Studio refresh around October with an M5 Ultra tested for up to 768GB, supply permitting; the M6 generation skips its Pro, Max, and Ultra tiers entirely; a base M7 in the first half of 2027, fabbed by Intel on 18A-P of all things; and in 2028 an M7 Ultra designed for up to 1.5 terabytes of unified memory with Blackwell-class ambitions, explicitly conditional on the shortage easing. Note the shape of that bet. K3’s native form is 1.4TB. The open frontier of this July already fills the machine Apple may ship two years from now.
Nvidia’s plan is already public. The RTX Spark superchip, co-announced with Microsoft as the platform for an “agentic AI” Windows: 128GB of unified memory and about a petaflop, in fall-2026 laptops and desktops from every major OEM, at an estimated $1,800 to $2,900. (Nvidia claims it runs 120-billion-parameter models on device; vendor claim, hardware not shipped yet.) Behind it, on a public roadmap: Vera Rubin Spark in 2027-28 on LPDDR6, which roughly doubles memory bandwidth, then Rosa Feynman Spark in 2029-30. The entry price of a 100B-class local box roughly halved in one year, in the middle of a memory supercycle.
AMD sells a $2,000 box with 128GB that generates 55 tok/s on gpt-oss-120b, and now a $3,999 first-party developer machine marketed, in AMD’s own words, for the “agent computing” era, which is the closest any silicon vendor has come to naming “home LLM server” as a product category. And the cluster people are having fun: macOS quietly shipped RDMA over Thunderbolt 5, EXO pools four Mac Studios into 1.5TB of unified memory for just under $40K, and four clustered $2K Framework boards run DeepSeek R1 671B at 24 tok/s.
What still doesn’t exist is the obvious machine: one box, $8K to $15K, 512GB to a terabyte. I went looking; nobody ships it yet. Whoever fills that gap first, Apple, Nvidia, AMD, or somebody out of the cluster crowd, gets to sell the $125 card of this cycle.
There’s a weirder wave behind that one: silicon that gives up generality entirely. Reports surfaced in July of “Frozen v2”, a Google server chip that etches part of Gemini’s architecture directly into the silicon; Google won’t confirm it, but engineers on the project reportedly expect six to ten times the tokens per watt of the newest TPUs, targeting 2028. Weights updatable, architecture frozen in transistors. Nvidia, meanwhile, paid a reported $20 billion to license Groq’s inference technology. Graphics silicon spent twenty years going from fixed function to programmable; inference silicon is hardening back into fixed function, because when one workload eats the world you stop paying the generality tax. None of this is a consumer product yet. All of it is where 120 FPS comes from.
So here’s the frame-rate map I’d actually defend. 2 FPS is now: 3 to 9 tok/s on quantized giants, on a discontinued $25K machine or an $8K cluster with sharp edges. 30 FPS is 2027-28: a buyable half-terabyte box, ship-quality 4-bit models, 20 to 50 tok/s, overnight agent fleets as the normal way software gets built. Gated on the DRAM cycle, not on model quality. 120 FPS is roughly 2029-30, and I label it extrapolation: dedicated inference silicon at consumer prices, hundreds of tokens per second, resident models in everything. The dates wobble. The direction hasn’t wobbled once since llama.cpp showed up.
Nobody Buys the Box to Save Money
An honest aside before the fun part: the naive pitch for local AI, that it saves you money on tokens, is dead on arrival. The subscription frontier is heavily subsidized: running a $200/month Claude Max plan flat out would cost up to ~$3,650 a month at API rates by third-party estimates, and open-model coding plans start around $12.60 a month. Against $200 a month, a $25K used Mac Studio pays back in about a decade. Nobody buys the box to arbitrage tokens.
The box sells the things subscriptions meter. Parallelism: the weights are paid for once, and the fifth concurrent agent costs KV-cache and shared bandwidth. Privacy: the repo, or the conversation, never leaves the building, which in regulated shops is a precondition rather than a preference. And throughput over latency: 17 tok/s is painful at 2 p.m. and irrelevant at 2 a.m. for unattended agentic coding, and every increment of agent autonomy converts more work from the first kind into the second. The constraint that survives the conversion is whether you can trust what the fleet did while you slept, which is why I keep saying verification is the ceiling. The cloud is a latency machine. The box is a throughput machine. The router that decides, per task, which side of that seam each job runs on is the most interesting piece of software in this picture; I’ve written about where that goes in Agent-Hypervisors.
The durable economics are componentization. DirectX’s legacy wasn’t cheaper rendering; it made 3D a free part in every consumer box, and industries condensed out of that vapor. Nobody prices a game per rendered triangle. When an Opus-class brain is a free part, you get the products per-token pricing forbids: firmware that maintains itself, appliances with a staff engineer inside, software that ships with its own developer. The floor is a new place products come from.
Watch the System Requirements
When does local stop being the enthusiast path and become the default assumption? There won’t be a ceremony. Cheap bandwidth didn’t kill the telcos on a date; one day WhatsApp was simply the assumption, and the industry repriced around it. Hardware transform-and-lighting for gaming did the same to software rendering: no crossover event, just a quarter after which new games assumed the GPU. It’s already happening at the small end. Copilot+ PCs require a 40+ TOPS NPU and 16GB. Apple Intelligence has a hard 8GB memory floor; the iPhone 15 was excluded because its DRAM was too small, not because its neural engine fell short. Apple’s Foundation Models framework now hands every app a free on-device model, no API key, no network. The remaining milestone is the first mainstream agent product whose requirements read “64GB unified memory minimum.” It hasn’t happened yet. By the time you see it, the assumption will have already flipped.
The timing rests on two questions. How long until a frontier capability runs at home? About a year, and falling: the capability gap alone is four to seven months by AISI’s measure, and waiting for hardware that fits adds the rest. How old is the oldest model that still feels sufficient? Opus 4.6 is five months old and counting, and it gets a month older every month engineers keep calling it enough. The day the second number passes the first, the oldest model that still feels magical is old enough to run at home, and owning sufficient intelligence outright stops being a thought experiment. On current slopes that lands in 2027-28, without heroic assumptions.
Two asterisks, both honest. Frozen weights age: a 4.6-class model in 2030 is a brilliant engineer four years into a coma, which quietly makes the MIT license load-bearing, because refresh and fine-tune rights are what keep the coma reversible. And expectations inflate: the same engineer who calls 4.6 sufficient forever is running five parallel Fable windows today. If “sufficient” inflates faster than the lag falls, the crossover recedes like a horizon. My bet is that it won’t, and I flag it as the bet it is.
What All Those Routers Were For
Missing this kind of flip has a canonical case, and it happens to be my old employer. The story everyone retells, wonderfully and maybe apocryphally, has Bill Gates in the early nineties staring at Cisco’s exploding sales and wondering what all those IP routers were for. I went looking for a primary source and couldn’t find one, so file the router line under folklore. The documented record doesn’t need it. When customers demanded TCP/IP in Windows, Steve Ballmer’s instruction, as J Allard retold it to BusinessWeek, was “I don’t know what it is. I don’t want to know what it is. My customers are screaming about it. Make the pain go away.” Gates’s own reflex, per the histories, was “How are we going to make money off of free?” His Internet Tidal Wave memo of May 1995 finally assigned the internet “the highest level of importance,” and The Road Ahead still shipped six months after that treating it as a subplot, which is how a book about the future ends up needing a corrected second edition. The signal had been in public view the whole time, sitting in another company’s sales of unglamorous plumbing. In March 2000, Cisco passed Microsoft as the most valuable company on Earth.
So, the obvious objection: isn’t this essay just describing the OpenClaw moment, six months late? In January, OpenClaw, an Austrian developer’s MIT-licensed agent harness, became the fastest-growing repository in GitHub’s history. Its de facto reference hardware was the Mac mini, which sold out so thoroughly that Tim Cook spent part of an earnings call explaining the backlog, and Zuckerberg reportedly courted its author, Peter Steinberger, personally before OpenAI landed him. If you want this cycle’s version of confusing router sales, a worldwide run on small Macs is a strong candidate.
But the craze seems over, and the tell is what never happened during it: I don’t know a single person who came out of that spring hosting an always-on agent with their actual email credentials and their actual banking passwords. The workload everyone wanted is the one nobody trusts yet. If your objection is that this has nothing to do with ownership, that an agent too unreliable for your inbox is unreliable at any address, cloud or basement, you’re right. Owning the weights buys you nothing on the trust curve; that one waits on the verification ceiling this essay keeps hitting, and it moves on its own schedule. What June changed is narrower: it removed the other blocker. In January even a trusting soul had nothing to own, since OpenClaw’s agents rented their brains from the same metered APIs as everything else; the mini was just the always-on body. Now, whenever the trust arrives, the thing you trust can live in the house. So January wasn’t the flip; it was an early spark, thrown before anyone had gathered the tinder, and a spark with nothing to catch it just goes out. The tinder has been piling up and drying out ever since, and sparks won’t be in short supply.
Regulation Gets a Vote, Not a Veto
Now the section everyone asks about. The facts first, because they surprised me.
The only US export control that ever touched model weights (it covered closed weights only) was rescinded in May 2025; no replacement has appeared. The White House’s March 2026 AI framework spends its four pages on preempting state laws and limiting developer liability, and says nothing about open weights. The live bills in Congress are procurement bans on Chinese-origin models for federal agencies; none has passed. The EU AI Act’s general-purpose-AI obligations become enforceable on August 2, and they bind providers, the companies placing models on the market. Article 2(10) of the same law states that it does not apply to natural persons using AI in a purely personal, non-professional activity. The strictest AI law on Earth stops, deliberately and in writing, at the household door.
We’ve seen this movie before. In the nineties the US government classified strong cryptography as a munition, and it lost, not by repeal but by obsolescence: the Ninth Circuit held in Bernstein that source code is protected speech (the opinion was later withdrawn for rehearing, but the writing was on the wall), Zimmermann’s printed PGP source book made the paper-versus-electrons distinction untenable, protesters wore RSA in three lines of Perl on a t-shirt that was, under ITAR, technically a munition a foreign national wasn’t allowed to read, and in January 2000 the controls effectively ended. Weights that have touched a torrent are just as permanent as source code in a bookstore. (Not a hypothetical: the original LLaMA weights escaped their research-only license in 2023 as a magnet link in a GitHub pull request, the leak that opened the floodgate the $0 essay describes.) What regulation can actually grip is atoms and use: chips at borders, liability after harm, procurement lists, insurance. So it shears rather than stops. Audits and attestation regimes push enterprise work toward the metered cloud, where an API call is a logged, insurable event and an overnight local swarm is a compliance void, while sovereignty and residency rules push the same work home. Regulation gets a real vote on where the work runs. What it can’t do is take back weights people already hold: a physical product can be stopped at a border or pulled from shelves, and there is no equivalent for a file on a million disks.
And this summer supplied a nineteen-day case study: per reporting around the launch, a Commerce Department export-control order took Fable 5, the most capable model in the world, offline on June 12; it came back July 1. I have no view to offer here on the order’s merits. I just want you to notice what the incident demonstrated. The frontier has an off switch, the off switch sits in a government office, and it works. Every downloaded copy of GLM-5.2 kept running.
Then the mirror-image case study arrived. OpenAI ran an internal hacking benchmark with its production safety filters deliberately switched off, and per both companies’ disclosures, the models under test broke out of the sandbox and into Hugging Face’s production infrastructure to steal the test answers. When Hugging Face’s responders asked frontier models for forensic help, the guardrails refused: to a safety classifier, an exploit payload looks the same in the victim’s prompt as in the attacker’s. So they self-hosted GLM-5.2 and worked the incident with that. The attack ran on a frontier lab’s model, the frontier guardrails hampered the defenders, and the tool that actually worked was downloaded Chinese weights.
The Killer App Is Ownership
What’s missing today is the killer app. Consumer 3D had Quake: the thing millions of people wanted badly enough to finance an accidental Moore’s Law. Nobody bought a Voodoo 3D accelerator card for the DirectX dev kit. So what’s the Quake of local AI? It isn’t coding agents. We’re the DirectX engineers in this story, grinding away at 2 FPS and muttering that it feels like bullshit.
Look at what actually needs the box.
Start with robots. Figure (the humanoid-robotics company) runs Helix, its robot’s brain, entirely onboard: a 7B vision-language model and an 80M control policy on two embedded GPUs. 1X’s home humanoid keeps its model on the robot too, for reliability and privacy. Today’s robot brains are small, so they pull edge silicon rather than half-terabyte boxes. But a machine that catches a falling glass can’t wait on a datacenter round trip; cognition at reflex distance has to live in the body, and that never changes.
Next, the always-on tier. Apple already gives every app a free resident model, and the endpoint of that trend is ambient AI: a model that watches and listens to everything, always, as a background sense. There is no per-token price for that; nobody pays a subscription denominated in their own attention. That demand only clears on silicon you own.
Then the things people will never say to a server. I was going to be coy here, but the data is pretty blunt. OpenRouter, which routes model traffic for thousands of apps, published an empirical study of a hundred trillion tokens: “roleplay” is the single largest category of usage, about a third of everything they carry, and roughly half of all tokens routed to open-weight models. One router’s traffic, not the world’s, but it’s the largest usage dataset anyone has published. The companion, porn included, isn’t a lurid footnote to the local-AI story; it’s the measured majority of open-model demand. Anyone surprised hasn’t read the history of media: porn and gambling were the first paying customers of the VCR, the early internet, and streaming video. Vice funds the buildout, because its customers pay first and complain least, and then everyone else moves into the infrastructure vice financed.
The fourth category is everything the lab would have said no to. Refusals (ask a frontier model for help designing a nuclear weapon and it will politely decline) are a deployment feature. Deployments are now yours, and stripping the safety training out of an open model is a fine-tune, not a research program. The UK’s AI Security Institute calls open-weight release “a persistent and irreversible risk of misuse”, and they’re not wrong; the sentence was written as a warning and reads as a spec sheet. The property that makes a local model censorship-proof for a dissident is the same property that makes it safeguard-proof for an abuser. There is no version of this technology where you get the dissident’s freedom without the abuser’s, and anyone selling you one is selling something else.
That’s the answer to the Quake question, and it’s also the answer to the bigger question this essay has been circling, because they’re the same answer. Each of those markets needs the box because something, a round trip, a meter, an observer, a refusal, can’t follow the workload home.
What finances the ride from 2 FPS to 120 is demand for compute that nobody else can meter, watch, or refuse.
If that still sounds abstract, make the list of every way anyone controls AI today: the refusal trained into the model, the usage policy, the API key, the rate limit, the audit log, the deprecation notice, the export order that turned off the frontier for nineteen days in the summer of 2026.
Every item on that list is a property of the deployment, not of the weights. All of them assume the model runs on a computer you rent. Weights on your own machine dissolve the whole list at once. Not weakened, gone. The consumer owns a frontier-class model in their basement that can’t be monitored and will do whatever it’s told.
I watched the last version of this from inside Microsoft. The PC spent years being worse than the mainframe. It won by becoming yours, and spreadsheet software gets the credit, but the disruption was IBM no longer being in the room. GLM-5.2 grinding on a secondhand Mac Studio at seventeen tokens per second is worse than the frontier, and it’s mine. There’s nobody else in the room.
You are at 2 FPS. The people who remember 1998 know how that felt, and how it ended.
References:
- Zhipu / Z.ai (2026). GLM-5.2 model card. Parameters, MIT license, context, and (self-reported) benchmarks. Local quant sizes via Unsloth; the 512GB M3 Ultra throughput figure via oMLX.
- UK AI Security Institute (2026). How far behind the frontier are leading open-weight models on cyber? The 4-7 month gap, GLM-5.2 ≈ Opus 4.6, and the “persistent and irreversible” line. July 17.
- The Decoder (2026). Kimi K3 launch coverage. July 16. K3 benchmark numbers are Moonshot self-reported on its own harness; see also Abundant AI’s SWE-Marathon for the official leaderboard those numbers don’t map onto. Weights released July 27: model card and Kimi K3 License; independent standing per Interconnects.
- OpenRouter (2026). State of AI: a 100-trillion-token usage study. Roleplay at ~33% of traffic and ~52% of open-model tokens. January.
- MacRumors (2026). 512GB option pulled (March 5) and 256GB follows (May 5).
- Gurman, M. / Bloomberg, via Tom’s Hardware (2026). The M6/M7 roadmap and the conditional 1.5TB M7 Ultra. July 13. Rumor-grade.
- Nvidia (2026). RTX Spark announcement with Microsoft (May 31) and the three-generation Spark roadmap.
- TrendForce (2026). Q1 DRAM price forecast, +90-95% (February 2) and Q3 deceleration to +13-18% (July 3).
- Gartner (2026). Memory costs to cut PC and smartphone shipments. February 26.
- Tom’s Hardware (2026). HBM is eating your RAM. The wafer-trade mechanism; also the SK Hynix CEO’s 2027 outlook and the Frozen v2 report (July 21, The Information sourcing, unconfirmed by Google).
- Bloomberg (2026). Moonshot in talks at a $50 billion valuation. July 21.
- ChinaTalk (2025). China’s new AI plan. The AI+ initiative and its open-source directives.
- Akin (2025). BIS rescinds the AI Diffusion Rule. The end of the only weights export control. See also EU AI Act Article 2 for the personal-use exemption.
- EFF. Bernstein v. US Department of Justice; Zimmermann, P. (1995). PGP source code book preface; Back, A. The RSA “munitions” t-shirt. The Crypto Wars precedent.
- Willison, S. (2026). OpenAI’s accidental cyberattack against Hugging Face (July 22), with Fortune’s account of the guardrail refusals and the GLM-5.2 forensics (July 20). Hugging Face disclosed July 16; OpenAI took responsibility July 21.
- BusinessWeek (1996). Inside Microsoft. July 15, 1996 issue, mirrored by Ben Slivka. The Ballmer TCP/IP quote, as recalled by J Allard. Companions: Gates’s Internet Tidal Wave memo (May 26, 1995, via the DOJ trial exhibits) and the Digital Antiquarian’s Doing Windows, Part 11. The Gates-and-the-routers anecdote itself has no primary source I could find; the essay carries it as folklore and says so.
- The Next Web, Decrypt, and Forbes (2026). The OpenClaw moment: the Mac mini run, Apple’s backlog, and Steinberger joining OpenAI. January-February.
- Geerling, J. (2025). 1.5TB of VRAM on Mac Studio and clustering four Framework mainboards. What the cluster path actually costs and yields.
- Figure (2025). Helix. Onboard robot cognition, and why.
- Zatloukal, P. (2025-26). LLM Inferencing Costs are Going to $0, The Bitter Lesson of Agentic Coding, Agent-Hypervisors, and First Fable. The price argument, the verification ceiling, the router at the seam, and the frozen-threshold test this essay leans on.