The Compiler Whisperer

A 125-Billion-Parameter Model on Two Ordinary PCs: The Build, the Numbers, and the Failure

By The Compiler Whisperer · October 8, 2026 · Part 2 of a two-part series

Part 1 answered whether a local model is worth running at all for legacy-code work. It was — with a 12B model doing in 72 seconds what a cloud API would have done for pennies, and the real dividing line being confidentiality, not cost. This article is the build itself: what I installed, on what hardware, what it actually measures, and the failure it gave me — which turned out to be the most transferable lesson of the two articles.

One note before anything else: I describe my machines by class and spec only. No machine names, no network addresses, no location. Where a config example needs an address, I use 192.0.2.10 — an address reserved by the IETF for documentation, not my network.

The shape is one you already know

If you built client–server applications in the 1990s — a thin front end on the desktop, the real work on a box down the hall — you already understand this setup. It is the same two-tier shape, with the roles renamed:

┌─────────────────────┐        LAN        ┌──────────────────────────┐
│  Box B — the agent  │ ────────────────▶ │  Box A — the inference   │
│  runs the agent,    │   HTTP POST to    │  server                  │
│  the scripts, the   │   /v1/chat/...    │  125B model, quantized,  │
│  editor, the browser│                   │  OpenAI-compatible API   │
└─────────────────────┘                   └──────────────────────────┘

Box B is deliberately weak. Box A does the arithmetic. The interface between them is a plain HTTP API — the modern descendant of the ODBC connection string and the RPC call: a service on a port, and a client that knows the address.

What I built

Box A — inferenceBox B — agent
CPU13th-gen Core i7Ryzen 5 7640HS (laptop-class)
RAM64 GB DDR416 GB
GraphicsRadeon RX 9060 XT, 16 GB VRAMintegrated only
Jobrun the model, serve the APIrun the agent and everything else

Box A is the "slightly generous" PC — the one you build when you enjoy building PCs. Box B is the machine an engineer uses for their day job. Neither is a workstation, and neither cost four figures.

The software is Strata (open source, MIT license), which runs Qwen3.8-Flash-Next — a 125-billion-parameter mixture-of-experts model — on consumer hardware. The trick, if you want the 1990s translation: the model is a team of 24,576 small specialists, and any given word only needs 10 of them. The busiest few thousand live in VRAM, all of them fit in 64 GB of RAM, and a 29 GB lookup table waits on the SSD. It is a virtual-memory strategy wearing a modern costume — anyone who tuned a database's buffer pool in 1995 already gets it.

Install: what actually happens

  1. Check the floor. Strata's published requirements: a 12 GB+ GPU (the RX 9060 XT is on the supported list) and 32 GB+ RAM. With 64 GB of RAM, the installer's recommended size for general work is the IQ3_XXS quantization — about 47 GB of model across RAM and VRAM. My Box A runs it with room to spare.
  2. Run the installer. On Windows: download, unzip, double-click START-HERE.bat. It inspects the GPU, RAM and disk, recommends a model size, and asks how much context to configure. I chose 120K (it rounds to a 131,072-token window).
  3. Wait for the download. This is the step nobody warns you about enough: the model files are 66–76 GB, plus a ~6 GB helper layer on first start. On a normal home connection this is an overnight job, and it needs real free disk — an SSD for the lookup table is not optional.
  4. Start the server. A start script launches the inference server, which listens on 127.0.0.1:8080 and speaks the OpenAI-compatible API — the same dialect every agent framework, IDE assistant and 30-line Python script already understands.
  5. Point Box B at Box A. On the agent box, the model endpoint is the inference box's LAN address instead of localhost — in documentation form: http://192.0.2.10:8080/v1. That is the entire "network architecture."

No client software is installed on Box B for the model itself. The agent just makes HTTP calls. If you ever maintained a two-tier app where the client broke because the server service wasn't started, you already know this failure mode too — and you already know how to fix it.

Speed, measured on my own box

I measured with a streaming request: send a prompt of known size, time the first token, time the last. Three runs, warm server, October 2026, IQ3_XXS quantization, 120K context configured. Raw JSON saved on disk.

Prompt sizeFirst token appears afterReading speed (prefill)Writing speed (decode)
65 tokens (one question)1.2 s—48 tok/s
~11K tokens (a long document)13 s821 tok/s45 tok/s
~36K tokens (a codebase)42 s859 tok/s44 tok/s
~125K tokens (a whole session)159 s788 tok/s40 tok/s

Two things to read out of that table.

Writing is fast enough to feel human. 40–48 tokens per second is faster than you read. Strata's published figures for a faster card (RX 9070 XT) are 52–60 tok/s, and the community benchmark table has an RX 9060 XT row at 15 tok/s baseline rising to 29 tok/s with speculative decoding enabled — my numbers sit above that row, consistent with Strata's built-in "guess-then-check" helper running. Different stacks, different methods; the honest summary is that a 16 GB Radeon writes answers at a readable speed and the spread between setups is real.

Reading is the bottleneck, and it scales linearly. A 125K-token prompt takes two and a half minutes just to start answering. Nothing about the model is "frozen" in that time — it is reading. Hold that thought; it is the exact mechanism that broke my agent.

The failure: the agent that could not forget

Here is the part I did not plan to write about. I had been working with the agent in one long session — the same conversation, hours of tool calls, file edits, measurements. Session history grows; agent frameworks handle that by compressing: the framework sends the whole history back through the model and asks it to write a summary, then continues with the summary instead of the raw transcript.

At some point the session grew past a threshold, the framework tried to compress, and compression failed — repeatedly. The signatures, from the agent's own logs:

Meanwhile Box A — the generous PC — stopped responding. It was not crashed. It was doing one 70,000-token read after another, at 100% utilization, and every retry queued another read behind it. To the person at Box B, a busy machine and a dead machine look identical. Every old-timer has seen this exact scene with a database server; the interface is different, the physiology is the same.

Why it broke — one misconfigured number

The post-mortem was embarrassingly simple. The framework's compression trigger was configured at 256,000 tokens. The model's context window at the time was 65,536. The trigger was four times larger than the window.

  intended:   history hits threshold ──▶ compress early, session stays healthy
  configured: history hits window limit ──▶ compress NOW, with a history
              the summarizing model can no longer fit in its own window

Compression is not free — it is a full re-read of the history (prefill, the slow direction from the table above) plus a summary generation. A history already past the window cannot be read by the model that must summarize it. The request queues, times out, gets retried, and each retry adds another doomed read to the inference box's queue. The system was not broken by size. It was broken by a threshold set above the limit it was supposed to protect.

I set the trigger to 60,000 tokens — about 45% of the current model's 131,072-token window. With my measured prefill speed (~800 tok/s), a 60K compression read costs about 75 seconds: inside the timeout, with margin, and the session never approaches the hard limit. The failure mode is gone because the compression now happens while it is cheap.

Two options I know about and have deliberately not enabled yet, for honesty: a dedicated small model to write the summaries (so the big model never spends its expensive reading capacity on housekeeping), and a hard session-length policy. If the early-trigger setting ever proves insufficient, those are the next two moves, in that order.

The fix that mattered more than the setting

While the compression was failing, I did the thing this site has argued about for ten articles: I wrote the session's state to a file — what was built, what was verified, what was pending, the exact next step — and started a new session that reads that file first.

The new session continued the work. Not approximately: it resumed mid-task, with the measured numbers, the file paths, and the open questions intact. The broken session was survivable because the knowledge was not only in it.

If that sentence gave you a flicker of recognition — The "Last Admin" Problem is the same story with a human instead of a session. A system whose knowledge lives in one head (or one conversation) dies when that container breaks. The prescription is identical: get the state onto disk, in a form someone — or something — else can pick up cold. An agent session is just the newest kind of last admin.

Risks and honest limits

RiskRealityMitigation
The download66–76 GB of model files; a real disk and a real connection are prerequisites.Check free space before installing; overnight download.
Long-context latency125K-token reads take minutes on a 16 GB card. Agents re-read history every turn, so a long session gets slower every turn.Keep sessions shorter; compress early; start new sessions deliberately.
Compression failureAny framework whose compression threshold exceeds the model window will fail exactly as mine did.Set the threshold below the window (I use ~45%); verify with the framework's status command.
Machine "freezes"Heavy prefill looks identical to a crash from the client side.Check the server's own logs before rebooting anything — nine times out of ten it is working, not dead.
Project churnStrata is young (created September 2026, MIT-licensed, active issue tracker). Commands and numbers move.Pin a release; re-read the docs after updates; treat published benchmarks as dated.

FAQ

Do I need a gaming PC?

You need a 12 GB+ GPU and 32 GB+ RAM — that is a gaming-class machine, but a mid-range one, not a halo one. The agent box needs neither: mine is a 16 GB laptop-class PC with no discrete graphics.

Is 120K context worth configuring?

It is worth having available and worth not living in. The window lets the agent hold a whole codebase or a whole session — but every turn pays the reading cost for what is in the window. A 60K session answers in tens of seconds; a 125K session starts its answers two and a half minutes later. Configure the big window, then keep sessions deliberately short.

Why two PCs instead of one?

Because the agent box stays responsive while the inference box grinds. On one machine, a long model read competes with your editor, browser and terminal for the same cores and memory. Two boxes is the oldest load-isolation trick in the book and it costs nothing if you already own both.

Can I rent the GPU instead of buying it?

Yes — hourly GPU hosting exists, and for occasional experiments it beats a 66 GB download and a weekend of setup. I keep a managed-hosting link here (I earn a small commission if you sign up; my disclosure is here); it is an option I describe, not one I use daily, because my inference box already exists.

What if my agent framework has no compression settings?

Then the handoff file is your compression. One page — goal, state, verified facts, next step — written before the session gets long, read by the next session. It works on every framework, including none.

Summary

Do this today: whatever you are working on right now — a migration, a rewrite, an agent session — write the one-page handoff. Goal, current state, what is verified, exact next step. It costs ten minutes and it is the difference between an interruption and a restart.

Honesty notes. All speed numbers are my own measurements (October 2026, IQ3_XXS quantization, streaming API, warm server) on the hardware described; your numbers will differ, and quantization choice moves them. Strata's published figures and the community benchmark rows are quoted as published, with their methods, not as equivalents to mine. Hardware is described by class and spec only — no machine names, addresses or location appear in this article, and the 192.0.2.10 address is the IETF's documentation range. Model versions, prices and project details change; verify before relying on any of them.