What a record has to contain
The test is whether the record alone lets you explain the decision without access to the running system.
- The input as received
- Or a stable reference to it. If the source can change, store a hash so you can prove what you saw.
- The retrieved context
- Which documents, which chunks, which versions. A retrieval answer is unexplainable without this.
- Model and prompt versions
- Both. "GPT-class model, some prompt" explains nothing. This is the field that turns a mystery into a diff.
- The output as produced
- Before any post-processing, and the processed form too if they differ.
- Confidence and alternatives
- Where the model provides them. A decision at 0.51 tells a different story from one at 0.99.
- Tools called
- With arguments and results, in order.
- The human action
- Approved, overridden, ignored, and by whom. The single most valuable field in the store.
Overrides are the point
If you record only what the system decided, you have a log. If you record what a person did about it, you have a feedback loop.
- Every override is a labelled eval case, arriving free, from the people whose judgement you are reproducing.
- A rising override rate on one category is your earliest signal that something has drifted, well before any aggregate metric moves.
- An override rate near zero means either the system is excellent or nobody is really reviewing. Both are worth knowing and they look identical from the outside.
- Overrides clustered on one operator usually mean a training or specification problem, not a model problem.
Ask for the override reason, from a short fixed list plus free text. Free text alone will not be filled in; a list alone will not tell you the interesting cases.
What it costs, and how to keep it affordable
Full records are large, mostly because of retrieved context. This is manageable if you decide it early rather than after storage becomes a line item.
- Store references to documents, not copies, with a hash to prove the version.
- Tier by age: full detail hot for ninety days, then compress and move to cold storage. Regulated sectors here will have a retention floor; find it before designing.
- Never let the audit write block the decision path. Write asynchronously, but write durably; a dropped record is invisible until you need it.
- Redact at write time, not at read time. Personal data that should not be in the store is much harder to remove once it is.
Make it answerable by the people who ask
An audit trail only a data engineer can query is not one. The people who ask are operations leads, compliance, and occasionally a regulator.
- One screen: enter a reference, see the decision, the reasoning, the sources and the human action.
- Plain language, not field names. "Matched to vendor Al Noor Trading because the tax number on page 2 matched" is the sentence somebody needs.
- Exportable, because a regulator will ask for a file rather than a login.
- Searchable by the identifiers the business uses, not by your internal ids.
Want us to run this with you?
The Audit is this method pointed at your systems, with a costed build plan at the end of it.
Schedule call
