ML Measurement Checklist

An ML project without a measurement framework has no acceptance criteria. This checklist covers the seventeen things that need to be in place before a model can be judged: ground truth, scoring, the deployment gate, retraining, and the people involved. Each item is explained in a measurement framework for machine learning projects.

How does the checklist work?

Seventeen questions in five groups, about ten minutes. Answer yes, uncertain or no; each answer shows guidance and a link to further reading. Only yes answers score.

The questions

Ground truth

1. Have you identified the decision points the model will automate or support?

Yes

Each decision point is a place to collect examples, which is where the ground truth dataset starts.

Uncertain

Walk through the process from the customer's side and mark every place where a person makes the same judgement repeatedly. Those are the candidates.

No

Start here; every later item on this list depends on it. A process walk is the usual way to find them.

2. Can you produce labelled examples of correct output for those decisions?

Yes

Then the task can be specified formally; the labelled set is the specification.

Uncertain

Look at what the organisation already records. Support tickets, rework queues and known-issue lists hold the input, the output and the correction in one place.

No

Without examples of correct output there is nothing to measure a model against. Producing them is the first piece of work, and it needs the domain expert.

3. Do you have at least 50 labelled examples, or a plan to produce them before evaluation begins?

Yes

Fifty well-chosen examples evaluate a model more reliably than a thousand ambiguous ones.

Uncertain

Count what exists before planning more. Historical corrections often supply the first few dozen.

No

Budget the labelling before the model. The cost is the number of examples, divided by examples per day, times the labeller's daily rate, and labelling is frequently the largest single cost in the project.

4. Do your examples cover failure cases as well as successes?

Yes

Then the evaluation can catch a regression on the cases that matter.

Uncertain

Production data leans towards the cases the system was built for. Check for empty fields, duplicates and inputs at the extremes.

No

Construct synthetic examples for the edge cases deliberately; production logs will not supply them.

Scoring

5. Have you defined what correct output means for each example, in terms you can count?

Yes

Then scoring is a comparison, and anyone can rerun it.

Uncertain

Where two experts would disagree on an example, the definition is not yet precise enough. Resolve the disagreements; they are part of the specification.

No

Write the definition down for each example before choosing a metric; the metric follows from the definition.

6. Have you assigned risk weights to failure classes, based on business consequence?

Yes

Then the weighted score measures what matters to the business.

Uncertain

Group the failures you have seen (routing, classification, extraction, boundary errors) and ask the domain expert which cost most.

No

Without weights, a model that scores 93% overall but fails on high-risk cases looks better than one that scores 89% evenly.

7. Have you set a weighted score threshold the model must clear before deployment?

Yes

That threshold is the acceptance criterion; sign-off becomes a check against it.

Uncertain

Set it from the business consequence of an error, before any results exist.

No

Without one, sign-off comes down to whether the team feels confident, which is where ML projects tend to stall.

8. Have you set a separate threshold for high-risk failure classes?

Yes

The second threshold catches a model that clears the overall score while failing on the cases that matter most.

Uncertain

Identify the failure classes with the highest business impact first; the threshold follows from them.

No

A single overall threshold can hide a model that fails on the most expensive cases. Add a second condition for those.

Deployment

9. Is the deployment gate defined before evaluation begins?

Yes

Then the results cannot move the goalposts.

Uncertain

Write the thresholds down and date them before the first evaluation run.

No

A gate set after the results are known tends to be set wherever the results landed. Fix it first.

10. Have you identified who owns the go/no-go decision?

Yes

One named owner keeps the decision from drifting between teams.

Uncertain

Name one person. A committee can advise, but the decision needs an owner.

No

Name the owner before the model is built; a gate is only useful if someone is accountable for applying it.

11. Do you have a plan for monitoring model performance after deployment?

Yes

A scheduled evaluation against the ground truth set will show whether the score is holding.

Uncertain

The simplest plan is to rerun the evaluation on a schedule and compare the score with the previous run.

No

ML systems fail quietly: the outputs become gradually less right as the inputs change. Monitoring is how you find out.

Retraining

12. Do you know how you will detect performance degradation in production?

Yes

Then drift shows up as a number before it shows up as a complaint.

Uncertain

Sample production outputs, have the domain expert correct them, and score the model against the corrections.

No

Degradation is usually gradual, which makes it easy to miss. Decide how you will measure it before it happens.

13. Have you estimated the cost of a retraining cycle?

Yes

Then retraining is a budget decision you can make on its merits.

Uncertain

Labels drive most of the cost: the number needed, divided by labels per day, times the labeller's daily rate.

No

Estimate it now. The formula needs three numbers, and it settles whether an improvement is worth paying for.

14. Have you identified who will produce new labels, and at what rate?

Yes

The rate sets how quickly the model can improve.

Uncertain

Labelling usually needs domain knowledge. Check whether the people who have it can spare the time.

No

Without a labeller, the model cannot improve after deployment. Name the person and measure their rate on a small batch.

Organisation

15. Do you have a domain expert who can confirm what correct output looks like?

Yes

They own two thirds of the problem: what the data means, and what correct output looks like.

Uncertain

Look for whoever resolves escalations today; they already apply the definition of correct.

No

Without one, nobody can say whether the model is right. This is the gap to close first.

16. Is that expert allocated time to the project, beyond occasional consultation?

Yes

Allocated time keeps labelling and review on schedule.

Uncertain

Agree a fixed number of hours a week and put it in their plan, or the work will slip behind their day job.

No

An expert consulted occasionally becomes the bottleneck. Budget their time as part of the project cost.

17. Do you have a plan for what happens when the model is wrong?

Yes

Then errors have a route: to a person, to a queue, or to a correction that feeds the labelled set.

Uncertain

Decide where a wrong output goes, who sees it, and whether the correction is recorded.

No

Every model is wrong some of the time. Plan the fallback (human review, a queue, a safe default) before deployment.

Your results

Scoring runs in your browser and needs JavaScript. The printable checklist below lists every question and answer.

Can I print the ML measurement checklist?

Yes. The version below lists every question with the points for each answer.

  1. Have you identified the decision points the model will automate or support?
    • [ ] Yes (1)
    • [ ] Uncertain (0)
    • [ ] No (0)
  2. Can you produce labelled examples of correct output for those decisions?
    • [ ] Yes (1)
    • [ ] Uncertain (0)
    • [ ] No (0)
  3. Do you have at least 50 labelled examples, or a plan to produce them before evaluation begins?
    • [ ] Yes (1)
    • [ ] Uncertain (0)
    • [ ] No (0)
  4. Do your examples cover failure cases as well as successes?
    • [ ] Yes (1)
    • [ ] Uncertain (0)
    • [ ] No (0)
  5. Have you defined what correct output means for each example, in terms you can count?
    • [ ] Yes (1)
    • [ ] Uncertain (0)
    • [ ] No (0)
  6. Have you assigned risk weights to failure classes, based on business consequence?
    • [ ] Yes (1)
    • [ ] Uncertain (0)
    • [ ] No (0)
  7. Have you set a weighted score threshold the model must clear before deployment?
    • [ ] Yes (1)
    • [ ] Uncertain (0)
    • [ ] No (0)
  8. Have you set a separate threshold for high-risk failure classes?
    • [ ] Yes (1)
    • [ ] Uncertain (0)
    • [ ] No (0)
  9. Is the deployment gate defined before evaluation begins?
    • [ ] Yes (1)
    • [ ] Uncertain (0)
    • [ ] No (0)
  10. Have you identified who owns the go/no-go decision?
    • [ ] Yes (1)
    • [ ] Uncertain (0)
    • [ ] No (0)
  11. Do you have a plan for monitoring model performance after deployment?
    • [ ] Yes (1)
    • [ ] Uncertain (0)
    • [ ] No (0)
  12. Do you know how you will detect performance degradation in production?
    • [ ] Yes (1)
    • [ ] Uncertain (0)
    • [ ] No (0)
  13. Have you estimated the cost of a retraining cycle?
    • [ ] Yes (1)
    • [ ] Uncertain (0)
    • [ ] No (0)
  14. Have you identified who will produce new labels, and at what rate?
    • [ ] Yes (1)
    • [ ] Uncertain (0)
    • [ ] No (0)
  15. Do you have a domain expert who can confirm what correct output looks like?
    • [ ] Yes (1)
    • [ ] Uncertain (0)
    • [ ] No (0)
  16. Is that expert allocated time to the project, beyond occasional consultation?
    • [ ] Yes (1)
    • [ ] Uncertain (0)
    • [ ] No (0)
  17. Do you have a plan for what happens when the model is wrong?
    • [ ] Yes (1)
    • [ ] Uncertain (0)
    • [ ] No (0)

Score out of 17: 0-11, measurement first; 12-16, nearly ready; 17-17, ready to evaluate.

www.bayis.co.uk/checklists/ml-measurement.html

Frequently asked questions

What is an ML measurement framework?

The set of things that decide whether a model is working: a ground truth dataset, a scoring function, risk weights for different failures, a deployment gate, and a plan for detecting and fixing degradation after launch.

How many labelled examples do we need?

Fewer than most teams expect. Fifty well-chosen examples with clear correct outputs evaluate a model more reliably than a thousand ambiguous ones, provided they include failure cases.

Why does an uncertain answer score the same as no?

Either way, the item is not yet in place. The difference matters for planning: uncertain items are usually the quickest to settle, which makes them the first deliverables.

Does this apply to projects built on large language models?

Yes. A new prompt, a new model or a revised pipeline is judged the same way: run it against the labelled set and check that it clears the gate.