An ML project without a measurement framework has no acceptance criteria. This checklist covers the seventeen things that need to be in place before a model can be judged: ground truth, scoring, the deployment gate, retraining, and the people involved. Each item is explained in a measurement framework for machine learning projects.
How does the checklist work?
Seventeen questions in five groups, about ten minutes. Answer yes, uncertain or no; each answer shows guidance and a link to further reading. Only yes answers score.
The questions
Ground truth
1. Have you identified the decision points the model will automate or support?
Yes
Each decision point is a place to collect examples, which is where the ground truth dataset starts.
Uncertain
Walk through the process from the customer's side and mark every place where a person makes the same judgement repeatedly. Those are the candidates.
Look at what the organisation already records. Support tickets, rework queues and known-issue lists hold the input, the output and the correction in one place.
Without examples of correct output there is nothing to measure a model against. Producing them is the first piece of work, and it needs the domain expert.
Budget the labelling before the model. The cost is the number of examples, divided by examples per day, times the labeller's daily rate, and labelling is frequently the largest single cost in the project.
5. Have you defined what correct output means for each example, in terms you can count?
Yes
Then scoring is a comparison, and anyone can rerun it.
Uncertain
Where two experts would disagree on an example, the definition is not yet precise enough. Resolve the disagreements; they are part of the specification.
Scoring runs in your browser and needs JavaScript. The printable checklist below lists every question and answer.
What next?
Three options, in order of commitment. (What we collect is set out in the privacy notice.)
Keep reading
The articles linked from your answers:
Can I print the ML measurement checklist?
Yes. The version below lists every question with the points for each answer.
Have you identified the decision points the model will automate or support?
[ ] Yes (1)
[ ] Uncertain (0)
[ ] No (0)
Can you produce labelled examples of correct output for those decisions?
[ ] Yes (1)
[ ] Uncertain (0)
[ ] No (0)
Do you have at least 50 labelled examples, or a plan to produce them before evaluation begins?
[ ] Yes (1)
[ ] Uncertain (0)
[ ] No (0)
Do your examples cover failure cases as well as successes?
[ ] Yes (1)
[ ] Uncertain (0)
[ ] No (0)
Have you defined what correct output means for each example, in terms you can count?
[ ] Yes (1)
[ ] Uncertain (0)
[ ] No (0)
Have you assigned risk weights to failure classes, based on business consequence?
[ ] Yes (1)
[ ] Uncertain (0)
[ ] No (0)
Have you set a weighted score threshold the model must clear before deployment?
[ ] Yes (1)
[ ] Uncertain (0)
[ ] No (0)
Have you set a separate threshold for high-risk failure classes?
[ ] Yes (1)
[ ] Uncertain (0)
[ ] No (0)
Is the deployment gate defined before evaluation begins?
[ ] Yes (1)
[ ] Uncertain (0)
[ ] No (0)
Have you identified who owns the go/no-go decision?
[ ] Yes (1)
[ ] Uncertain (0)
[ ] No (0)
Do you have a plan for monitoring model performance after deployment?
[ ] Yes (1)
[ ] Uncertain (0)
[ ] No (0)
Do you know how you will detect performance degradation in production?
[ ] Yes (1)
[ ] Uncertain (0)
[ ] No (0)
Have you estimated the cost of a retraining cycle?
[ ] Yes (1)
[ ] Uncertain (0)
[ ] No (0)
Have you identified who will produce new labels, and at what rate?
[ ] Yes (1)
[ ] Uncertain (0)
[ ] No (0)
Do you have a domain expert who can confirm what correct output looks like?
[ ] Yes (1)
[ ] Uncertain (0)
[ ] No (0)
Is that expert allocated time to the project, beyond occasional consultation?
[ ] Yes (1)
[ ] Uncertain (0)
[ ] No (0)
Do you have a plan for what happens when the model is wrong?
[ ] Yes (1)
[ ] Uncertain (0)
[ ] No (0)
Score out of 17:
0-11, measurement first; 12-16, nearly ready; 17-17, ready to evaluate.
www.bayis.co.uk/checklists/ml-measurement.html
Frequently asked questions
What is an ML measurement framework?
The set of things that decide whether a model is working: a ground truth dataset, a scoring function, risk weights for different failures, a deployment gate, and a plan for detecting and fixing degradation after launch.
How many labelled examples do we need?
Fewer than most teams expect. Fifty well-chosen examples with clear correct outputs evaluate a model more reliably than a thousand ambiguous ones, provided they include failure cases.
Why does an uncertain answer score the same as no?
Either way, the item is not yet in place. The difference matters for planning: uncertain items are usually the quickest to settle, which makes them the first deliverables.
Does this apply to projects built on large language models?
Yes. A new prompt, a new model or a revised pipeline is judged the same way: run it against the labelled set and check that it clears the gate.
left to answer.
All questions answered. The reading list below collects the articles linked from your answers.