Skip to content
ENع
Schedule call
All guides
Governance7 min read

Rollback and kill switches

At some point the system will start doing something wrong at volume, and the person who notices will not be the person who built it. Everything about that moment should be decided in advance, because it will not be a good time to design anything.

What the switch has to reach

A kill switch that only stops new work leaves a queue draining into production for the next twenty minutes.

  1. Stop accepting new work at ingress, and say so, rather than silently dropping it.
  2. Stop workers picking up queued work, without losing the queue. The work must still be there afterwards.
  3. Decide what happens to work in flight: finish the current item or abandon it. Finishing is usually right if writes are idempotent, and abandoning is right if the problem is the writes.
  4. Above all, stop the write-back. Reading and reasoning can continue safely; committing cannot.

Test it, in production, on a quiet afternoon, before you need it. An untested kill switch is a belief.

Degrade rather than stop, where you can

A binary switch is a blunt instrument, and the cost of using it discourages using it early. Intermediate positions get used sooner, which is the point.

Gate everything
Keep running but route every decision to human approval. Throughput drops, correctness is preserved. This is the right first move for a quality problem.
Fall back a model
For a provider incident. Slower or slightly less accurate beats stopped.
Shed the low-value work
Keep the critical path, pause the batch. Protects the thing that matters under capacity pressure.
Read only
Answers and summaries continue, writes stop. Useful when the problem is downstream.

Rolling back a prompt is not rolling back code

Reverting a deployment restores behaviour. Reverting a prompt restores the instructions, which is not the same thing, because the model may have changed underneath you.

  1. Pin the model version explicitly. If you are on a floating alias, your behaviour can change with no deployment at all, and no rollback will restore it.
  2. Version prompt and model together, and record the pair on every decision. Rolling back means restoring the pair.
  3. Re-run the eval set after a rollback. A rollback is a change, and it deserves the same gate as any other.
  4. Records written under the bad version are still wrong. Rollback stops the bleeding; the backfill is a separate piece of work and should be scoped separately.

Who can pull it

The most common failure is not a missing switch, it is that the only person allowed to use it is asleep.

  1. The operations lead on shift can trigger it, without an engineer. If it needs a deployment, it is not a kill switch.
  2. Write the criteria down, so it is not a judgement call at 2am. "More than N exceptions in ten minutes" or "any write above X that was not gated".
  3. Make it obvious and slightly hard: a button, with a confirmation, that also opens an incident and notifies the channel.
  4. Nobody is ever criticised for using it. The moment somebody is, it stops being used, and you have lost the control entirely.

Want us to run this with you?

The Audit is this method pointed at your systems, with a costed build plan at the end of it.

Schedule call
Tell us the number you want to move.Schedule call