Why accuracy misleads
If 3% of your invoices have a problem, a model that says "no problem" every time is 97% accurate and completely useless. Every real workflow has this shape somewhere.
The second issue is asymmetry. Missing a fraudulent transaction and flagging a legitimate one are both errors and they are not remotely the same error. One costs the loss, the other costs a phone call. A metric that weights them equally is optimising the wrong thing by construction.
Pick from the cost, not from the textbook
Ask which mistake actually hurts, then choose accordingly.
- Missing one is expensive
- Optimise recall on the class that matters. Fraud, safety, compliance breaches. You will accept false positives and staff for them.
- A false alarm is expensive
- Optimise precision. Anything that interrupts a customer or a senior person. Alert fatigue is a real cost and it kills adoption.
- Both matter
- Use a weighted score with the weights set by the business, and write down who set them.
- It is a ranking
- Precision at k, because nobody reads past the first screen.
- It is generative
- A rubric scored by a second model, validated against human scores on a sample. Not similarity to a reference answer, which measures phrasing.
The metric we ask for first
Before any of the above: what fraction of cases can be closed without a person, at an error rate the business accepts?
That single number is what the project is for. It maps directly to cost, it forces the error rate to be stated rather than assumed, and it makes the trade explicit: raise the automation rate and errors rise, tighten the threshold and more work goes to people. Everything else is diagnostic.
Have someone name the acceptable error rate before you start, and get it in writing. A team that will not commit to a number in advance will reject any number you produce.
Report it so it cannot be gamed
- Always split routine from exception. The aggregate hides the only interesting part.
- Report the confusion matrix, not just the headline. Which way it is wrong matters more than how often.
- Include the abstention rate. A model that declines 30% of cases is not 100% accurate on the rest in any useful sense.
- Show the trend across versions. A single number tells you nothing about whether you are improving.
Want us to run this with you?
The Audit is this method pointed at your systems, with a costed build plan at the end of it.
Schedule call
