Skip to content
ENع
Schedule call
All docs
Step-by-stepStart hereUpdated Aug 2026 · 18 min read

Ship your first AI workflow

Six steps, in order, with what to produce at the end of each one. This is the sequence we run inside an Audit and Build, written so your team can run it without us.

6
weeks end to end
1
workflow, not a department
2
people from your side
10
readiness questions to clear
Gulls and swans over open water at a jetty
Steps
Before you start1 · Pick the workflow2 · Map it as it really runs3 · Build the eval set4 · Prove a thin slice5 · Integrate and gate6 · Roll out and hand overAppendix · Readiness checkRun it with us

Before you start

Two people from your side make or break this: one engineer with production access, and one operator who does the work today. Roughly a fifth of their time for six weeks. If you cannot free those two, fix that before anything else. No amount of engineering compensates for guessing at the process.

  1. 01Confirm the two people, in writing, with their managers.
  2. 02Agree the one number that should move, and where it is measured today.
  3. 03Get read access to a representative slice of real data, not a sanitised export.
  4. 04Decide where the system will run: your cloud tenancy, on-premise, or air-gapped.
Pick a workflow the operator complains about. Enthusiasm from the person doing the work is worth more than executive sponsorship at this stage.
STEP 1

Pick the workflow

Half a day

One workflow, with a start and an end you can point at. Not “customer service”, but the specific path from an inbound message to a resolved case in the system of record.

Score your candidates on four axes and take the highest total. Volume and repetition make the payoff real; a clean data source and a writable target system make it buildable.

Scoring worksheet
workflow            volume  repetition  data  writable  total
------------------  ------  ----------  ----  --------  -----
invoice intake         5         5        4       4       18
supplier onboarding    3         4        2       3       12
quote generation       4         3        3       4       14

# 1 = low, 5 = high. Take the highest total, not the loudest request.
Two hands playing an upright piano
A candidate workflow scored with the operator in the room, before any technical scoping.
If your top candidate scores below 12, you are not ready to build. That finding is worth more than a pilot that stalls in month four.
STEP 2

Map it as it really runs

2 to 3 days

Sit with the operator and watch a real case travel the whole path. Write down every step, every system touched, and, most importantly, every exception they handle without thinking about it.

You are looking for four things: where information gets re-typed, where a decision needs judgement, where an approval happens, and what the operator does when the case is unusual. That last list is where pilots die.

  1. 01Shadow three real cases end to end. Do not interview. Watch.
  2. 02Record each step as: input → decision → output → system it lands in.
  3. 03List the exceptions. Ask “what do you do when…” until they run out of answers.
  4. 04Mark every point where money, a customer, or a regulated record is touched. Those become approval gates.
A mountain ridge under a wide sky
The process map that comes out of step two: steps, systems, exceptions and approval points on one page.
Expect the documented process and the real process to differ. The real one is the specification.
STEP 3

Build the eval set

2 days

Before writing the system, assemble a scored set of cases from your own history. Around a hundred decided cases is usually enough to be meaningful, and it must include the exceptions from step two.

This is the single highest-leverage artefact in the project. It tells you whether the system works, catches regressions on every change, and, because it comes from your own records, settles arguments that a vendor benchmark cannot.

evals/cases.jsonl
{"id":"INV-2291","input":{"source":"email","attachments":1},
 "expected":{"vendor":"Al Noor Trading","total":4820.00,"gl":"6100"},
 "tags":["standard"]}
{"id":"INV-2317","input":{"source":"email","attachments":3},
 "expected":{"vendor":"Gulf Logistics","total":1195.50,"gl":"6220"},
 "tags":["multi-page","exception:split-invoice"]}
  1. 01Pull 100 recently decided cases, weighted toward the messy ones.
  2. 02Record the decision a good human made, not what the model should say.
  3. 03Tag exceptions explicitly so you can score them separately.
  4. 04Define the pass bar against the current human baseline. Measure that baseline first; it is rarely as high as people assume.
If your team cannot rerun the eval set without you, the handover has already failed. Keep it in the repository, runnable with one command.
STEP 4

Prove a thin slice

Week 1 to 2

Build the narrowest version that touches real data end to end. Ugly is fine. What matters is that a genuine record travels the full path inside two weeks: in from the source, decided, and out to a place a human can see.

Run it against the eval set, then read the failures individually. The failure pattern tells you whether this is a prompt problem, a retrieval problem, a data problem, or a “this should not be AI” problem.

Running the eval
$ npm run eval -- --set evals/cases.jsonl

  standard        92/94   97.9%   ✓ above bar (95%)
  exceptions      11/19   57.9%   ✗ below bar (85%)
  ---------------------------------------------
  overall        103/113  91.2%   avg $0.04 / case

  ! 6 of 8 exception failures share tag: exception:split-invoice
A line of cypress trees along a ploughed field
A thin slice running against real records in week two, with the eval report beside it.
A clustered failure tag like the one above is good news. It is one rule to handle, not a general quality problem.
STEP 5

Integrate and gate

Week 3 to 4

Now make it land. The output goes into the system of record, idempotently, with a rollback path, and the approval gates from step two get built in rather than promised.

Security review runs alongside this step, not after it. Bring the reviewer the deployment topology, the audit log format and the access model while there is still time to change them.

  1. 01Write back with an idempotency key so a retry cannot create a duplicate.
  2. 02Log every model call, tool use and write with actor, cost and outcome.
  3. 03Put a human gate on anything touching money, a customer or a regulated record.
  4. 04Build the kill switch and confirm your team, not ours, can throw it.
Write-back with rollback
POST /invoices                        # idempotent
Idempotency-Key: INV-2291-v3

{ "vendor": "Al Noor Trading",
  "total": 4820.00,
  "gl": "6100",
  "source": "momentem/intake@1.4.2",
  "decision_id": "d_8f21c",           # links to the audit record
  "requires_approval": false }

# reversal: POST /invoices/INV-2291/void { "decision_id": "d_8f21c" }
A dense city skyline seen from above
The approval queue an operator actually uses, with the audit record behind each decision.
If an automated action cannot be undone in one step, it needs an approval gate instead of automation. No exceptions.
STEP 6

Roll out and hand over

Week 5 to 6

Go live on one team, on live work, with the switch in their hands. Widen it as the numbers justify, not as the roadmap demands.

Handover is a deliverable, not a meeting. It is done when your engineer ships a change to the system, on their own, and the eval set catches a regression they introduced deliberately.

  1. 01One team, live traffic, rollback available and rehearsed.
  2. 02Watch the number from step one weekly against its baseline.
  3. 03Write the runbook: alerts, common failures, what to do at 3am.
  4. 04Pair on-call until your engineer has handled an incident without calling us.
  5. 05Test the handover: have them break something on purpose and confirm the eval catches it.
Skyscrapers converging overhead against a pale sky
Week six: the number moving, and the runbook that lets your team own it.

At this point the system belongs to you in every meaningful sense: the code is in your repository, the evals are in your CI, and the person who fixes it at 3am works for you.

Appendix · Readiness check

Before step one, answer these with evidence rather than intent: a number someone owns, a system name, a job title. Anything you cannot answer is scope, and that list is usually more useful than the answers you already have.

Which single number should move, and what is it today?With an owner who agrees it is theirs.
How much of the work is the routine path?This sets the ceiling on any automation.
Where does the input actually live?System of record, not the reporting layer.
Is there decided history to score against?Twelve months is usually enough.
Which system must the output land in, and can it be written to?An import-only target changes the build.
Can an automated write be reversed in one action?If not, it needs an approval gate.
Who must approve what?Money, customers and regulated records, explicitly.
Can data leave your network?If not: your tenancy, on-premise, or air-gapped.
What must be logged, and retained how long?To your own policy, not ours.
Who owns the eval set after handover?If the answer is us, the engagement is not finished.
Want us to run it with you?

Same six steps, pointed at your systems, with senior engineers alongside your team and a build plan you could execute with anyone.

Schedule call

Related guides

Evaluation

Building an eval set from your history

10 min
Integrations

Idempotent write-back

9 min
Governance

Designing approval gates

8 min
Tell us the number you want to move.Schedule call