The seven
- Ingress
- Whatever brings work in: a queue, a webhook, a mailbox, a scheduled sweep. Owns deduplication and the idempotency key.
- Orchestration
- Decides what happens in what order and what to do when a step fails. Deterministic code, not a model.
- Model calls
- The steps that need a model, each behind an interface that takes your types.
- Tools
- Governed access to your systems. Read and write, each with its own permissions and its own audit line.
- Gates
- Where a person must approve before an action commits. Explicit, configurable, and logged either way.
- Decision store
- What was decided, on what input, by which model and prompt version, with what confidence.
- Evals
- The scored set that runs on every change, on your own decided cases.
Only two of those are AI. The other five are ordinary systems engineering, which is why AI projects staffed entirely with data scientists stall at the demo.
Orchestration is code, not a model
The most common expensive mistake is letting a model decide control flow. It is seductive because it demos beautifully and it fails in ways you cannot reproduce.
- A model choosing between three tools will choose differently on the same input next Tuesday. Your incident review will have nothing to point at.
- Retries become unbounded. A loop with a model in the condition has no guaranteed exit.
- Cost becomes unpredictable, because the number of calls is now an emergent property.
- You cannot test it, because the state machine does not exist anywhere you can enumerate.
Use a model where judgement is genuinely required: reading a document, weighing an ambiguous case, drafting a summary. Use code for everything about sequencing, retries and branching. A state machine you can draw on a whiteboard is worth more than an agent you cannot.
The two everyone skips
The decision store and the eval set. Both feel like overhead in month one and both are the difference between a system you can operate and one you can only restart.
The decision store is what lets you answer "why did it do that" three weeks later, which is the first question anyone asks. It needs the input as received, the output as produced, the model and prompt versions, the confidence, the tools called, and whether a human overrode it. That last field is the most valuable data you will collect: overrides are a labelled training and eval set arriving for free.
The eval set is what lets you change anything. Without it, every model upgrade, prompt edit and dependency bump is a leap, so the team stops making them, and the system ossifies within a quarter.
If you build only one of the seven before the first workflow ships, build the decision store. You can add evals from it later. You cannot recover decisions you never wrote down.
Where it runs
In the Gulf this is usually settled before the architecture is. The shape holds either way, and the parts that move are fewer than people expect.
- Hosted models through your own cloud account: the default. Data stays in your tenancy, the provider does not train on it, and you keep the frontier.
- Self-hosted open weights in your VPC: for workloads that genuinely cannot egress. See the guide on self-hosting for what this actually costs.
- Split: sensitive extraction on-premise, general reasoning hosted, with a boundary you can point at in a review. This is the common answer for banking and government here, and it works.
Note that only the model-call component moves. Ingress, orchestration, tools, gates, the decision store and evals are identical in all three, which is why designing around the boundary rather than the provider keeps the option open.
Want us to run this with you?
The Audit is this method pointed at your systems, with a costed build plan at the end of it.
Schedule call
