The Compiler Whisperer

Local LLM vs Cloud API for Legacy Code Analysis: What My Own Test Showed

By The Compiler Whisperer · October 7, 2026 · Part 1 of a two-part series

People who hear that I run a 27-billion-parameter language model on my own PC usually react the same way: "Isn't a cloud API cheaper? You must need a monster machine." That reaction is reasonable — and I decided to test it instead of arguing with it.

This article is the result. I built a small fictional Access/VBA system, ran it through three local models on my own hardware, scored the answers against a ground truth I wrote myself, and compared the numbers against published cloud API prices. Everything here is a measurement I actually made, with the raw outputs kept on disk. Where I estimate, I say so.

The test design

Two constraints shaped the experiment:

The prompt was deliberately boring and identical for every model: "Extract every business rule in this VBA code. For each rule give: (1) plain-language statement, (2) code location, (3) confidence, (4) questions you would ask the original developer. Do not invent anything not in the code."

The hardware question, answered with my own numbers

My setup is two ordinary home PCs on a LAN, and neither is a workstation. One runs a local inference server that exposes an OpenAI-compatible API; the other runs the agent and drives that server from a 30-line Python script. The agent box is a Ryzen 5 7640HS with 16 GB of RAM and integrated graphics — the kind of PC an engineer buys for themselves and keeps for six years.

On the inference box, a 12-billion-parameter model quantized to 4-bit (about 7 GB on disk) produced a complete, usable answer in 72 seconds. A 27-billion model also runs there — slowly, and as you will see, not successfully. So the first half of the "you need a monster PC" claim is simply false for this class of task: an ordinary home PC was enough.

What the inference box has not shown me is speed on big contexts. That is a real limit, and I'll price it honestly below.

The results

The 12B model: 12 of 12 rules, zero inventions, 72 seconds

The 12B instruction-tuned model (Gemma-family, 4-bit quantized) returned a complete inventory of 15 items in 72 seconds. It found all 12 ground-truth rules, each attributed to the correct function. Its line numbers drifted a few lines (it cited the tax multiplication as line 55; it is line 50) — close enough to navigate, not close enough to trust blindly.

Two behaviors impressed me more than the hit rate:

The 27B model: correct analysis, no deliverable

My main model — the 27B one people assume is the crown jewel — failed the same task.

The run generated 4,000 tokens in 128 seconds, and every one of them went into the model's private reasoning trace. The answer field came back empty. The saved output is 16,240 characters of raw deliberation ("We need answer the user's request…") truncated mid-sentence before any document was emitted.

Here is the uncomfortable part: reading the deliberation trace, the 27B model had found the rules correctly — internal transfers, the three discount multipliers, zone D, tax ordering, priority levels, the Sunday and year-end date shifts, the ticket cap, the two tolerance constants, the credit limit and its hardcoded manager check. Its analysis was arguably the best of the three. It just never wrote it down. A reasoning model that cannot close its answer is useless in a pipeline, no matter how good its thinking is.

The 8B model: fast, overconfident, unstable

The 8B model (a DeepSeek-distilled Qwen) produced 22 numbered items — all 12 real rules present, but 10 of the 22 are padding (it listed tables and forms as "rules"), with line numbers wrong throughout (off by roughly 13 lines) and a habit of stamping "Missing Info: None" — 17 times — on rules whose business meaning it could not possibly know.

Re-run with a larger output budget, the same model at the same temperature produced only 10 rules — it dropped the priority rules and the tax rule entirely, and misstated the internal-transfer rule as "exempt from tax" when the code zeroes the whole total. Same model, same prompt, two different inventories. Run-to-run variance is a finding in itself.

The tacit knowledge: 0 of 7, every model

No model recovered any of the seven backstory items. The 12B model's contribution was to ask about them — "Why is zone D exempt? Is D always the company's own depot?" "Why hardcode the string 'MGR' when a user table exists?" — questions that are, verbatim, the questions I ask in real modernization projects. The backstory answers came from me, the fictional developer.

This matches what I have seen for thirty years: the code is the easy part. The knowledge lives in the head of the person who wrote it in 2003, or in whoever kept it running after they left.

Is the cloud actually cheaper? The honest arithmetic

The claim "cloud is cheaper because local needs expensive hardware" is half right, and the half that is wrong is usually the bigger half for a solo consultant. Here is the arithmetic as I work it.

Cloud side. Published API prices for frontier models run roughly $3 per million input tokens for a mid-tier model and $15 input / $75 output for the top tier (Claude Sonnet and Opus, as published in Anthropic's pricing page; the exact numbers move, so check the page before quoting them). A 116-line VBA file is about 1,500 tokens. One analysis pass — file plus instructions plus a long answer — is on the order of 5,000 tokens total. At Sonnet-class prices that is a couple of cents per file; at Opus-class prices, tens of cents. Even 100 files is single-digit dollars. For occasional analysis, the cloud is effectively free, and my PC cannot compete with that.

Local side. My PC already exists — I do my day job on it. The marginal cost of running models on it is zero dollars and some electricity. The real costs are different ones: a 12B model answers in a minute; a 70B model on ordinary PC RAM is slow to the point of irritation; and if I genuinely needed frontier-class quality locally, the hardware would be a four-figure purchase (a 24 GB GPU or a large-RAM Mac) plus the time to maintain the stack. That purchase only pays off if I run many thousands of analyses or if the data cannot legally leave my machine.

The variable that decides it is not price — it is confidentiality. Client code is confidential. That is not a preference, it is the condition of getting hired. For the work I actually do, uploading a client's Access database or VBA module to any third-party API is off the table regardless of price. For my own experiments, my own writing, and public sample code, the cloud is cheaper and better and I should use it.

So the received wisdom is correct when the data may leave the building. It is irrelevant when it may not. And there is a third option neither extreme covers: a cloud plan that runs on isolated infrastructure you rent by the hour — which is where managed GPU hosting comes in, priced per hour rather than per token, and worth knowing about even if you never buy it. (If you go looking, I earn a small commission on some sign-ups through links like this one; it changes nothing about my conclusions, and my full disclosure is here.)

What I would actually do, by situation

SituationMy recommendation
Occasional analysis of code you are allowed to send outCloud API. Cents per file, frontier quality, no maintenance.
Client code under a confidentiality obligationLocal only, and only models you can verify. Or no AI at all.
High-volume analysis (hundreds of modules, repeated passes)Local starts to win on cost — but measure your own throughput before buying hardware.
Exploring whether AI helps your modernization work at allStart local on the PC you already own, with a 12B model. Zero cost, real data.

The finding that matters more than the cost question

The cost debate is the question everyone asks. The experiment answered a more important one: the bottleneck is not the model, and it is not the money. A free-to-run 12B model extracted every rule from 116 lines of VBA in a minute. What it could not do — what no model did — was recover the seven decisions that only the developer knew. The tacit layer still requires a human who can read old code, recognize which questions matter, and sit with the person who kept the system alive.

AI changes the price of the mechanical half of modernization. It has not changed the price of the human half. That is the whole thesis of this site, and now I have a measurement to back it.

Part 2 of this series will walk through the local setup itself — the inference server, the quantized model, the 30-line script — on hardware equivalent to a normal office PC, so you can reproduce the experiment.

Honesty notes. The sample system is fictional and was written by me for this test; no client code was used at any point. The cloud prices above are from published pricing pages as of October 2026 and will change — verify before relying on them. The local run times are from my own two-PC home setup (quantized GGUF models on a local inference server) and will differ on yours.