Skip to content
ENع
Schedule call
All guides
Infrastructure7 min read

Cost controls that hold

AI spend surprises people because it is one of the few infrastructure costs that scales with how carelessly the software is written rather than with how many users it has. The good news is that three controls account for most of the reduction, and all three are cheap to add.

Find out where it goes first

Almost every team we have looked at was wrong about which step cost the most, usually by a factor of five. Measure before you optimise.

  1. Record input tokens, output tokens, cached tokens and model per call, tagged with the workflow step.
  2. Group by step, not by day. One step is normally 60% or more of the bill.
  3. Split input from output. Output is several times dearer, so a chatty step with a short input can dominate a step that reads a whole document.
  4. Look at the tail. A p99 that is twenty times the median usually means a retry loop nobody has noticed.

Prompt caching, which is most of the win

If your system prompt carries instructions, a schema, examples or retrieved context, you are paying full price to send the same tokens repeatedly. Caching that prefix is typically the single largest reduction available.

  1. Put everything stable at the front: tools, then instructions, then reference material. Put anything that varies at the very end.
  2. Mark the boundary at the end of the stable part. Marking it at the end of the whole prompt means every request writes a new entry and none is ever read.
  3. Check nothing varies inside the prefix. A timestamp, a request id or a reordered JSON key invalidates everything after it, and this is the most common reason caching appears not to work.
  4. In a conversation, mark the end of the most recent turn as well, so a long thread re-reads its own prefix instead of paying for it again.

A cache read is a fraction of the price of a fresh read, and a write is slightly more than one. If your traffic re-reads a prefix even a few times, it pays immediately.

Routing and ceilings

Routing: send the easy majority to a small model and escalate only what needs depth. You already have the signal for this if you have been storing confidence. In the pipelines we run, between 70% and 90% of cases never need the expensive model.

Ceilings are not an optimisation, they are a safety device, and every public endpoint needs them:

  1. A hard cap on tokens per conversation or per job, enforced in code, not in a prompt.
  2. A bound on tool-call rounds within one turn, so a loop cannot run indefinitely.
  3. Rate limits per caller, held somewhere shared rather than in the memory of one instance.
  4. A daily spend alarm that pages someone, at a level you would notice but survive.

Every one of those exists because of an incident somebody had. The loop bound in particular: an agent that can call a tool that produces work for itself will do so until something stops it.

Want us to run this with you?

The Audit is this method pointed at your systems, with a costed build plan at the end of it.

Schedule call
Tell us the number you want to move.Schedule call