Skip to content
ENع
Schedule call
All insights
Method9 min read

Running AI on-premise in the Gulf

Every enterprise conversation in this region opens with sovereignty. Most of them should not end where they begin, because the requirement as first stated is rarely the requirement, and the architecture that satisfies it is usually two steps cheaper than the one being proposed.

The requirement is narrower than the sentence

"The data cannot leave the country" is where these conversations start. An hour of questions usually turns it into something more specific and much more tractable: one class of field, under one regulation, with processing and storage treated separately, and de-identified data frequently out of scope entirely.

We have had this conversation perhaps thirty times. In most of them the binding requirement covered customer identifiers and account records, and the analytical content, the part a model actually needs to reason over, was not in scope once identifiers were removed. That single distinction is the difference between a data centre programme and a tokenisation layer.

Ask to read the clause. Not the policy, not the summary, the clause. In roughly a third of cases there is no clause and the requirement is an internal position, which can be discussed on its merits.

What self-hosting actually costs

The case for self-hosting is usually made on unit economics, comparing GPU rental to per-token pricing. That comparison leaves out most of the cost.

Utilisation
You pay for the card idle. A working-hours workload runs near 30%, which triples the real unit cost immediately.
Redundancy
One card is a demo. Two, across zones, is the production floor.
An owner
Somebody now runs inference: batching, memory, quantisation, upgrades, a rota. This is the largest line and it is never in the spreadsheet.
Standing still
Hosted models improve without you. Yours improves when someone runs a migration, and the gap compounds.

Our rule of thumb: self-hosting starts to compete on cost above roughly a million calls a month on a narrow, stable task at high utilisation. Below that it is a sovereignty decision rather than an economic one, and it deserves to be argued as such rather than dressed up as savings.

The split that usually wins

For most regulated clients here the answer is neither fully hosted nor fully on-premise. It is a boundary drawn where the sensitivity actually sits.

Extraction and classification over the sensitive corpus run inside the client tenancy, on open weights. Those tasks are narrow, and open models do them well. Anything requiring real reasoning depth runs on a hosted frontier model, over content that has already had its identifiers stripped or tokenised. The boundary is a documented interface with a data-flow statement attached, which is exactly the artefact an auditor wants and exactly what a purely on-premise build usually fails to produce.

It also preserves the option. When the next open model closes the gap you move more work behind the boundary. When it does not, you have not committed to a capability ceiling.

Where it genuinely is the only answer

  1. Air-gapped by mandate. Defence, some government functions, some critical infrastructure. There is no alternative and the cost question does not arise.
  2. A regulator who has already refused the hosted arrangement in writing. Argue it once, then build for it.
  3. Volume that genuinely clears the threshold on a stable narrow task, where the arithmetic works with the full cost included.

What is not on that list: general assistants, anything whose specification is still moving, and anything needing frontier reasoning. Those are precisely where the open-weight gap is widest and where you will feel it every week rather than once at procurement.

What we tell clients

Establish the requirement precisely before designing anything, work down the architecture list rather than up, and write the data-flow statement before you write code. If sovereignty is genuinely binding, build for it properly and price it honestly, including the engineer who will own inference.

The failure we see most is a sovereign build justified on cost that was never going to be cheaper, running a model two generations behind, maintained by nobody in particular. It satisfies the letter of the requirement and delivers very little, which is the worst of both positions.

Want us to run this with you?

The Audit is this method pointed at your systems, with a costed build plan at the end of it.

Schedule call
Tell us the number you want to move.Schedule call