Structure it so a stranger can find the part to change
A prompt that runs together into one wall of instruction cannot be edited safely, because no one can tell which sentence is load-bearing. Give it fixed sections in a fixed order.
- Role and scope
- What this call is for, and explicitly what it is not for. One short paragraph.
- Inputs
- What it will be given, in what shape. Name the fields.
- Rules
- The constraints, as a list. One rule per line, each independently removable.
- Output contract
- Exactly what to return. Prefer a schema and strict tool use over describing a format in prose.
- Examples
- Two or three, and at least one hard case. Last, because they are the part most often changed.
The order matters for caching as well as for reading: everything stable goes first, so a provider cache can be reused across calls, and only the varying tail is reprocessed.
Rules that survive contact
- Write each rule so it can be tested. "Be helpful" cannot fail. "Never return a total that is not present in the source document" can.
- Say what to do when the rule cannot be met. An instruction with no escape hatch produces invention, which is the failure mode you are trying to avoid.
- Prefer one specific rule to three general ones. General rules interact in ways nobody predicts.
- Delete rules that no longer earn their place. Prompts accumulate defensive instructions against failures that were fixed elsewhere months ago, and every one of them costs tokens and attention.
If you cannot say which eval case a rule was added for, that is a strong sign it can go. Remove it, run the set, and find out.
Version it like code, because it is code
A prompt is the logic of the step. It belongs in the repository, in review, with history.
- One file per prompt, in version control, not a string in a service or a row in a database an operator can edit unreviewed.
- Change it through a pull request, with the eval delta in the description. "Improves exception accuracy from 71% to 84%, no regression on routine" is a reviewable claim.
- Stamp the prompt version onto every decision record. When output quality moves you need to know what changed and when.
- Never edit a prompt directly in production. It is the single most common cause of a system that worked last week and nobody can explain.
The handover test
The real question is whether a second engineer can change this safely on their first week. Test it directly.
- Give the prompt and its eval set to somebody who did not write it.
- Ask them to make one specific behavioural change: tighten a rule, add a category, change what happens on an ambiguous input.
- Watch where they hesitate. Every hesitation is a place the structure failed, not a place they were slow.
- Have them run the evals and read the delta. If they cannot tell whether their change was an improvement, the eval set is the thing to fix.
A prompt that passes this is a component. One that does not is a person, and people leave.
Want us to run this with you?
The Audit is this method pointed at your systems, with a costed build plan at the end of it.
Schedule call
