Before you start
Two people from your side make or break this: one engineer with production access, and one operator who does the work today. Roughly a fifth of their time for six weeks. If you cannot free those two, fix that before anything else. No amount of engineering compensates for guessing at the process.
- 01Confirm the two people, in writing, with their managers.
- 02Agree the one number that should move, and where it is measured today.
- 03Get read access to a representative slice of real data, not a sanitised export.
- 04Decide where the system will run: your cloud tenancy, on-premise, or air-gapped.
Pick the workflow
One workflow, with a start and an end you can point at. Not “customer service”, but the specific path from an inbound message to a resolved case in the system of record.
Score your candidates on four axes and take the highest total. Volume and repetition make the payoff real; a clean data source and a writable target system make it buildable.
Map it as it really runs
Sit with the operator and watch a real case travel the whole path. Write down every step, every system touched, and, most importantly, every exception they handle without thinking about it.
You are looking for four things: where information gets re-typed, where a decision needs judgement, where an approval happens, and what the operator does when the case is unusual. That last list is where pilots die.
- 01Shadow three real cases end to end. Do not interview. Watch.
- 02Record each step as: input → decision → output → system it lands in.
- 03List the exceptions. Ask “what do you do when…” until they run out of answers.
- 04Mark every point where money, a customer, or a regulated record is touched. Those become approval gates.
Build the eval set
Before writing the system, assemble a scored set of cases from your own history. Around a hundred decided cases is usually enough to be meaningful, and it must include the exceptions from step two.
This is the single highest-leverage artefact in the project. It tells you whether the system works, catches regressions on every change, and, because it comes from your own records, settles arguments that a vendor benchmark cannot.
- 01Pull 100 recently decided cases, weighted toward the messy ones.
- 02Record the decision a good human made, not what the model should say.
- 03Tag exceptions explicitly so you can score them separately.
- 04Define the pass bar against the current human baseline. Measure that baseline first; it is rarely as high as people assume.
Prove a thin slice
Build the narrowest version that touches real data end to end. Ugly is fine. What matters is that a genuine record travels the full path inside two weeks: in from the source, decided, and out to a place a human can see.
Run it against the eval set, then read the failures individually. The failure pattern tells you whether this is a prompt problem, a retrieval problem, a data problem, or a “this should not be AI” problem.
Integrate and gate
Now make it land. The output goes into the system of record, idempotently, with a rollback path, and the approval gates from step two get built in rather than promised.
Security review runs alongside this step, not after it. Bring the reviewer the deployment topology, the audit log format and the access model while there is still time to change them.
- 01Write back with an idempotency key so a retry cannot create a duplicate.
- 02Log every model call, tool use and write with actor, cost and outcome.
- 03Put a human gate on anything touching money, a customer or a regulated record.
- 04Build the kill switch and confirm your team, not ours, can throw it.
Roll out and hand over
Go live on one team, on live work, with the switch in their hands. Widen it as the numbers justify, not as the roadmap demands.
Handover is a deliverable, not a meeting. It is done when your engineer ships a change to the system, on their own, and the eval set catches a regression they introduced deliberately.
- 01One team, live traffic, rollback available and rehearsed.
- 02Watch the number from step one weekly against its baseline.
- 03Write the runbook: alerts, common failures, what to do at 3am.
- 04Pair on-call until your engineer has handled an incident without calling us.
- 05Test the handover: have them break something on purpose and confirm the eval catches it.
At this point the system belongs to you in every meaningful sense: the code is in your repository, the evals are in your CI, and the person who fixes it at 3am works for you.
Appendix · Readiness check
Before step one, answer these with evidence rather than intent: a number someone owns, a system name, a job title. Anything you cannot answer is scope, and that list is usually more useful than the answers you already have.
Same six steps, pointed at your systems, with senior engineers alongside your team and a build plan you could execute with anyone.
Schedule call






