The arithmetic people do, and the one they should
The comparison is almost always GPU rental against per-token pricing, which makes self-hosting look cheap at any real volume. The line items that get left out are the ones that dominate.
- Left out: utilisation
- You pay for the card whether or not traffic is arriving. A workload with a working-hours shape runs at maybe 30% utilisation, which triples the true unit cost.
- Left out: redundancy
- One card is not a production deployment. Two, in two zones, is the floor.
- Left out: the engineer
- Somebody now owns inference: batching, memory, upgrades, an on-call rota. This is the largest line and it never appears in the spreadsheet.
- Left out: falling behind
- Hosted models improve without you doing anything. Yours improves when someone does a migration.
Our rule of thumb: self-hosting starts to compete on cost somewhere north of a million calls a month on a narrow task, at high utilisation. Below that, it is a sovereignty decision, not an economic one, and it should be argued as such.
When it genuinely is the answer
- The data legally cannot leave, and no hosted arrangement satisfies the regulator. This is the common case here, and it is a complete answer on its own.
- The workload is narrow, high volume and stable. Classification and extraction at millions of calls a month, on a task that has not changed in a year.
- Air-gapped by requirement. Defence, some government, some critical infrastructure. There is no alternative and the cost question does not arise.
Note what is not on that list: general assistants, anything needing frontier reasoning, and anything whose specification is still moving. Those are the workloads where the gap between open weights and the frontier is widest and where you will feel it every week.
What it takes to run properly
- Pick a serving stack with continuous batching. Serving one request at a time wastes most of the card, and this single choice moves throughput by several times.
- Size for the memory, not the parameter count: weights, plus the KV cache, which scales with concurrency and context length and is what actually runs you out.
- Quantise deliberately and measure the accuracy cost on your own eval set, not on a public benchmark. Sometimes it is free. Sometimes it removes exactly the capability you were relying on.
- Put a queue in front. Without one, a burst degrades every request instead of delaying some.
- Pin the model version and treat an upgrade as a deployment with an eval gate, because that is what it is.
The split that usually wins
For most regulated clients here the answer is not either. It is a boundary drawn where the sensitivity actually is.
Extraction and classification over the sensitive corpus run on-premise, because that is where the data is and those tasks are narrow enough for open weights to do well. Anything general, or anything needing depth, runs on a hosted frontier model with the sensitive fields already stripped or tokenised. The boundary is a documented interface, which is what a regulator wants to see.
It also keeps the option open. When the next open model closes the gap you move more behind the boundary; when it does not, you have not committed.
Want us to run this with you?
The Audit is this method pointed at your systems, with a costed build plan at the end of it.
Schedule call
