A 125-Billion-Parameter Model on Two Ordinary PCs: The Build, the Numbers, and the Failure
Part 1 answered whether a local model is worth running at all for legacy-code work. It was — with a 12B model doing in 72 seconds what a cloud API would have done for pennies, and the real dividing line being confidentiality, not cost. This article is the build itself: what I installed, on what hardware, what it actually measures, and the failure it gave me — which turned out to be the most transferable lesson of the two articles.
One note before anything else: I describe my machines by class and spec only. No machine names, no network addresses, no location. Where a config example needs an address, I use 192.0.2.10 — an address reserved by the IETF for documentation, not my network.
The shape is one you already know
If you built client–server applications in the 1990s — a thin front end on the desktop, the real work on a box down the hall — you already understand this setup. It is the same two-tier shape, with the roles renamed:
┌─────────────────────┐ LAN ┌──────────────────────────┐ │ Box B — the agent │ ────────────────▶ │ Box A — the inference │ │ runs the agent, │ HTTP POST to │ server │ │ the scripts, the │ /v1/chat/... │ 125B model, quantized, │ │ editor, the browser│ │ OpenAI-compatible API │ └─────────────────────┘ └──────────────────────────┘
Box B is deliberately weak. Box A does the arithmetic. The interface between them is a plain HTTP API — the modern descendant of the ODBC connection string and the RPC call: a service on a port, and a client that knows the address.
What I built
| Box A — inference | Box B — agent | |
|---|---|---|
| CPU | 13th-gen Core i7 | Ryzen 5 7640HS (laptop-class) |
| RAM | 64 GB DDR4 | 16 GB |
| Graphics | Radeon RX 9060 XT, 16 GB VRAM | integrated only |
| Job | run the model, serve the API | run the agent and everything else |
Box A is the "slightly generous" PC — the one you build when you enjoy building PCs. Box B is the machine an engineer uses for their day job. Neither is a workstation, and neither cost four figures.
The software is Strata (open source, MIT license), which runs Qwen3.8-Flash-Next — a 125-billion-parameter mixture-of-experts model — on consumer hardware. The trick, if you want the 1990s translation: the model is a team of 24,576 small specialists, and any given word only needs 10 of them. The busiest few thousand live in VRAM, all of them fit in 64 GB of RAM, and a 29 GB lookup table waits on the SSD. It is a virtual-memory strategy wearing a modern costume — anyone who tuned a database's buffer pool in 1995 already gets it.
Install: what actually happens
- Check the floor. Strata's published requirements: a 12 GB+ GPU (the RX 9060 XT is on the supported list) and 32 GB+ RAM. With 64 GB of RAM, the installer's recommended size for general work is the IQ3_XXS quantization — about 47 GB of model across RAM and VRAM. My Box A runs it with room to spare.
- Run the installer. On Windows: download, unzip, double-click
START-HERE.bat. It inspects the GPU, RAM and disk, recommends a model size, and asks how much context to configure. I chose 120K (it rounds to a 131,072-token window). - Wait for the download. This is the step nobody warns you about enough: the model files are 66–76 GB, plus a ~6 GB helper layer on first start. On a normal home connection this is an overnight job, and it needs real free disk — an SSD for the lookup table is not optional.
- Start the server. A start script launches the inference server, which listens on
127.0.0.1:8080and speaks the OpenAI-compatible API — the same dialect every agent framework, IDE assistant and 30-line Python script already understands. - Point Box B at Box A. On the agent box, the model endpoint is the inference box's LAN address instead of localhost — in documentation form:
http://192.0.2.10:8080/v1. That is the entire "network architecture."
No client software is installed on Box B for the model itself. The agent just makes HTTP calls. If you ever maintained a two-tier app where the client broke because the server service wasn't started, you already know this failure mode too — and you already know how to fix it.
Speed, measured on my own box
I measured with a streaming request: send a prompt of known size, time the first token, time the last. Three runs, warm server, October 2026, IQ3_XXS quantization, 120K context configured. Raw JSON saved on disk.
| Prompt size | First token appears after | Reading speed (prefill) | Writing speed (decode) |
|---|---|---|---|
| 65 tokens (one question) | 1.2 s | — | 48 tok/s |
| ~11K tokens (a long document) | 13 s | 821 tok/s | 45 tok/s |
| ~36K tokens (a codebase) | 42 s | 859 tok/s | 44 tok/s |
| ~125K tokens (a whole session) | 159 s | 788 tok/s | 40 tok/s |
Two things to read out of that table.
Writing is fast enough to feel human. 40–48 tokens per second is faster than you read. Strata's published figures for a faster card (RX 9070 XT) are 52–60 tok/s, and the community benchmark table has an RX 9060 XT row at 15 tok/s baseline rising to 29 tok/s with speculative decoding enabled — my numbers sit above that row, consistent with Strata's built-in "guess-then-check" helper running. Different stacks, different methods; the honest summary is that a 16 GB Radeon writes answers at a readable speed and the spread between setups is real.
Reading is the bottleneck, and it scales linearly. A 125K-token prompt takes two and a half minutes just to start answering. Nothing about the model is "frozen" in that time — it is reading. Hold that thought; it is the exact mechanism that broke my agent.
The failure: the agent that could not forget
Here is the part I did not plan to write about. I had been working with the agent in one long session — the same conversation, hours of tool calls, file edits, measurements. Session history grows; agent frameworks handle that by compressing: the framework sends the whole history back through the model and asks it to write a summary, then continues with the summary instead of the raw transcript.
At some point the session grew past a threshold, the framework tried to compress, and compression failed — repeatedly. The signatures, from the agent's own logs:
- a compression request that timed out — the summary call never completed within the framework's 300-second cap;
- the HTTP connection to the inference box being forcibly closed mid-request;
- and finally the blunt verdict: the session was too large to compress under the model's context window — the history no longer fit in the very model that had to read it.
Meanwhile Box A — the generous PC — stopped responding. It was not crashed. It was doing one 70,000-token read after another, at 100% utilization, and every retry queued another read behind it. To the person at Box B, a busy machine and a dead machine look identical. Every old-timer has seen this exact scene with a database server; the interface is different, the physiology is the same.
Why it broke — one misconfigured number
The post-mortem was embarrassingly simple. The framework's compression trigger was configured at 256,000 tokens. The model's context window at the time was 65,536. The trigger was four times larger than the window.
intended: history hits threshold ──▶ compress early, session stays healthy
configured: history hits window limit ──▶ compress NOW, with a history
the summarizing model can no longer fit in its own window
Compression is not free — it is a full re-read of the history (prefill, the slow direction from the table above) plus a summary generation. A history already past the window cannot be read by the model that must summarize it. The request queues, times out, gets retried, and each retry adds another doomed read to the inference box's queue. The system was not broken by size. It was broken by a threshold set above the limit it was supposed to protect.
I set the trigger to 60,000 tokens — about 45% of the current model's 131,072-token window. With my measured prefill speed (~800 tok/s), a 60K compression read costs about 75 seconds: inside the timeout, with margin, and the session never approaches the hard limit. The failure mode is gone because the compression now happens while it is cheap.
Two options I know about and have deliberately not enabled yet, for honesty: a dedicated small model to write the summaries (so the big model never spends its expensive reading capacity on housekeeping), and a hard session-length policy. If the early-trigger setting ever proves insufficient, those are the next two moves, in that order.
The fix that mattered more than the setting
While the compression was failing, I did the thing this site has argued about for ten articles: I wrote the session's state to a file — what was built, what was verified, what was pending, the exact next step — and started a new session that reads that file first.
The new session continued the work. Not approximately: it resumed mid-task, with the measured numbers, the file paths, and the open questions intact. The broken session was survivable because the knowledge was not only in it.
If that sentence gave you a flicker of recognition — The "Last Admin" Problem is the same story with a human instead of a session. A system whose knowledge lives in one head (or one conversation) dies when that container breaks. The prescription is identical: get the state onto disk, in a form someone — or something — else can pick up cold. An agent session is just the newest kind of last admin.
Risks and honest limits
| Risk | Reality | Mitigation |
|---|---|---|
| The download | 66–76 GB of model files; a real disk and a real connection are prerequisites. | Check free space before installing; overnight download. |
| Long-context latency | 125K-token reads take minutes on a 16 GB card. Agents re-read history every turn, so a long session gets slower every turn. | Keep sessions shorter; compress early; start new sessions deliberately. |
| Compression failure | Any framework whose compression threshold exceeds the model window will fail exactly as mine did. | Set the threshold below the window (I use ~45%); verify with the framework's status command. |
| Machine "freezes" | Heavy prefill looks identical to a crash from the client side. | Check the server's own logs before rebooting anything — nine times out of ten it is working, not dead. |
| Project churn | Strata is young (created September 2026, MIT-licensed, active issue tracker). Commands and numbers move. | Pin a release; re-read the docs after updates; treat published benchmarks as dated. |
FAQ
Do I need a gaming PC?
You need a 12 GB+ GPU and 32 GB+ RAM — that is a gaming-class machine, but a mid-range one, not a halo one. The agent box needs neither: mine is a 16 GB laptop-class PC with no discrete graphics.
Is 120K context worth configuring?
It is worth having available and worth not living in. The window lets the agent hold a whole codebase or a whole session — but every turn pays the reading cost for what is in the window. A 60K session answers in tens of seconds; a 125K session starts its answers two and a half minutes later. Configure the big window, then keep sessions deliberately short.
Why two PCs instead of one?
Because the agent box stays responsive while the inference box grinds. On one machine, a long model read competes with your editor, browser and terminal for the same cores and memory. Two boxes is the oldest load-isolation trick in the book and it costs nothing if you already own both.
Can I rent the GPU instead of buying it?
Yes — hourly GPU hosting exists, and for occasional experiments it beats a 66 GB download and a weekend of setup. I keep a managed-hosting link here (I earn a small commission if you sign up; my disclosure is here); it is an option I describe, not one I use daily, because my inference box already exists.
What if my agent framework has no compression settings?
Then the handoff file is your compression. One page — goal, state, verified facts, next step — written before the session gets long, read by the next session. It works on every framework, including none.
Summary
- A 125B model runs on a mid-range 16 GB Radeon and 64 GB of RAM, installed by double-clicking one script — the barrier is the 66–76 GB download and patience, not exotic hardware.
- Writing answers is fast (40–48 tok/s measured); reading long contexts is the real cost (125K tokens ≈ 159 seconds before the first word appears).
- The failure was a compression threshold set above the model's context window — a number, not a mystery. Set it to ~45% of the window and the failure mode disappears.
- The insurance that actually saved the work was a one-page handoff file on disk. Sessions, like last admins, are containers — and containers break.
Do this today: whatever you are working on right now — a migration, a rewrite, an agent session — write the one-page handoff. Goal, current state, what is verified, exact next step. It costs ten minutes and it is the difference between an interruption and a restart.
Honesty notes. All speed numbers are my own measurements (October 2026, IQ3_XXS quantization, streaming API, warm server) on the hardware described; your numbers will differ, and quantization choice moves them. Strata's published figures and the community benchmark rows are quoted as published, with their methods, not as equivalents to mine. Hardware is described by class and spec only — no machine names, addresses or location appear in this article, and the 192.0.2.10 address is the IETF's documentation range. Model versions, prices and project details change; verify before relying on any of them.