Skip to content
ENع
Schedule call
All guides
Infrastructure8 min read

Sizing from 0 to 1M+ users

Scaling an AI system is not like scaling a web service. The bottleneck is rarely your compute; it is a provider rate limit, a queue you did not have, or a cost curve that was linear right up until it was not. Here is what breaks, roughly in order.

Up to a thousand a day

Nothing breaks. This is the range where teams over-build and lose two months.

  1. Call the provider synchronously. You do not need a queue yet.
  2. Do not cache. You will cache the wrong thing and debug a stale answer for a day.
  3. Do build the decision store now, while the volume is small enough to read by eye. This is the single highest-value thing you can do early.
  4. Do build the eval set now, from the cases you are already deciding.

The mistake at this stage is architecture. The mistake at every later stage is not having built the two things above at this stage.

Ten thousand a day

The first real breakages, and they are all about time rather than volume.

  1. Provider rate limits, and they bite in bursts rather than on the daily average. A month-end run that is fine at 400 an hour dies at 4,000 in twenty minutes.
  2. Synchronous calls start timing out somewhere they are not retried, and work vanishes silently.
  3. A tail of slow requests holds connections open and starves everything else.

The fix is a queue and a worker pool, with retries that are idempotent. This is the point where the ingress component earns its place, and where the idempotency key you defined at the start saves you from double-posting invoices.

A hundred thousand a day

  1. Cost becomes the constraint before capacity does. At this volume a careless prompt is a five-figure monthly line item, and prompt caching stops being an optimisation and becomes the difference.
  2. Prompt caching pays properly here: keep everything stable at the front of the prompt and only the varying tail at the end, and the input bill falls by most of itself.
  3. Batch anything that is not interactive. Overnight reconciliation does not need a live call.
  4. The eval set becomes too slow to run in full on every commit. Split it: a fast subset on every push, the whole thing nightly and before release.
  5. A single provider becomes a real availability risk. This is when a second adapter is worth writing, not before.

A million and beyond

At this point the interesting problems are organisational as much as technical.

  1. Model choice is a fleet decision. Routing cheap cases to a small model and escalating the rest is worth a great deal, and needs the confidence score you have been storing since day one.
  2. Self-hosting starts to pencil out for the high-volume narrow steps, purely on unit cost. Check the arithmetic in the self-hosting guide before believing it.
  3. Drift becomes measurable and matters. Your input distribution moves, and the eval set that was representative last year is not.
  4. Regional capacity, not global capacity, is what limits you. Ask your provider about the region you are actually in.

What has not changed across four orders of magnitude: the orchestration is still deterministic code, and every decision is still written down.

Want us to run this with you?

The Audit is this method pointed at your systems, with a costed build plan at the end of it.

Schedule call
Tell us the number you want to move.Schedule call