Data Maturity Checklist

Data maturity is how far an organisation can trust, find and use its data in consistent, repeatable ways, and it decides what AI can do before any model is chosen. This checklist covers storage, third-party services, reproducibility, ownership, governance and use, following what is data maturity?

How does the checklist work?

Fourteen questions in six groups, about eight minutes. Answer yes, partly or no; each answer shows guidance and a link to further reading, and the result names the area to start with. We record which answers are chosen, without any name or contact details; those reach us only if you use one of the forms at the end.

The questions

Storage

1. Can anyone on the team say where the latest version of a key dataset lives?

Yes

A consistent location is the first sign of storage maturity. Retention rules and version history come next.

Sometimes

Pick the three datasets the organisation relies on most and record where the authoritative copy of each lives. That list is the start of a catalogue.

No

Start here. Until the current version can be found, every other question on this list is hard to answer.

2. Do reports draw on the data directly, without manual downloads?

Yes, from a shared store or dashboard

Then reports update with the data, and everyone reads the same numbers.

Some do; others are exported by hand

Each manual export is a copy that stops updating the moment it is made. Automate the one that is repeated most often.

Most are exported by hand

Reports built from exported copies drift from the source and from each other. A shared, queryable store is the fix.

3. Do you know how long data is kept, and can you see how it has changed?

Yes

Retention rules and an audit trail let you say what a record looked like last year, with confidence.

Partly

Write the retention rules down, then keep raw records unchanged so that later changes can be traced.

No

Keep raw records as they arrive and never overwrite them; the raw record is the audit trail.

External services

4. Do you know which third-party services hold your data?

Yes, and we keep a list

Then your dependencies are visible. Check that each service allows a full export.

Roughly

List them, with what each holds and who administers it. The list often turns up a service nobody remembers signing up for.

No

List them first. Data you cannot locate cannot be backed up, moved or joined to anything else.

5. Could you move a key service to a different provider?

Yes

Then changing provider is a planned project with a known cost.

With difficulty

Test a full export from the service you depend on most. What comes out, and in what shape, is the real measure.

No

Then the provider controls both the data and the terms. Plan an export route before you need it.

Reproducibility

6. Can you trace a key figure back to its source data and the steps that produced it?

Yes

Then a change in a trend reflects the data, with the method held constant.

For some figures

Start with the numbers the board sees. Record the source, the steps and the definitions behind each one.

No

Without a trace, a change in a number could be a change in the world or a change in the method. Record how each key number is produced.

7. Can you repeat a data collection (a survey, a cohort, a metric) the same way next time?

Yes

Then collection is part of a pipeline, and comparisons over time hold.

Partly

Write down the collection steps and the definitions they use, and keep them with the data.

No

Data treated as the one-off result of someone's effort is hard to compare or scale. Script the collection you repeat most.

Ownership

8. If the person who calculates a key number left tomorrow, could someone else produce it?

Yes

Then the knowledge lives in the process, where it survives staff changes.

With some effort

Pair someone with the current owner on the next run, and write each step down as they go.

No

A common single point of failure in small teams. Capture the steps now, while the person is still there.

9. Do different teams get the same answer from the same data?

Yes

Then shared definitions are doing their job.

Usually

Where the answers differ, the definitions differ. Agree one definition per metric and keep it next to the data.

No

Metrics open to interpretation turn meetings into arguments about numbers. Define each key metric once, in one place.

10. Are data definitions kept close to the data, in the schema, metadata or a readme?

Yes

Then anyone can read what a field means without asking.

Some of them

Start with the fields people ask about most often.

No

Definitions held in people's heads leave with them. Document the schema where the data lives.

Governance

11. Do access permissions reflect roles and regulatory requirements?

Yes

Then access can grow with the organisation without growing the risk.

Partly

Review who can see personal and confidential data first, and remove access nobody needs.

No

Start with personal and confidential data: who can see it, and why. UK GDPR expects access to be limited to what each purpose needs.

12. Would you notice if data went missing or was overwritten?

Yes, we have alerts

Then data problems surface as alerts before they surface in a decision.

Eventually

Add checks to the feeds behind key reports: row counts, arrival times, and values out of range.

No

Silent loss is found only when something breaks. Start with an alert on missing or late data for the most important feed.

Use

13. Can people get the data they need without asking someone to pull it?

Yes

Self-service access means the data is organised for use.

Some of them

Note which requests recur, and publish those as a standing report or dataset.

No

Every request that needs a person creates a queue and another copy. Publish the requests that recur.

14. Is a new metric easy to add?

Yes

Then the data model has room for new questions, and the history you have accumulated may answer more of them than the current reports show.

It takes a project

A new metric should be a query. When it takes a project, the data usually lacks a shared model.

No

Then the data answers only the questions it was set up for. A shared, well-defined data model is what changes that.

Your results

Scoring runs in your browser and needs JavaScript. The printable checklist below lists every question and answer.

Can I print the data maturity checklist?

Yes. The version below lists every question with the points for each answer.

  1. Can anyone on the team say where the latest version of a key dataset lives?
    • [ ] Yes (2)
    • [ ] Sometimes (1)
    • [ ] No (0)
  2. Do reports draw on the data directly, without manual downloads?
    • [ ] Yes, from a shared store or dashboard (2)
    • [ ] Some do; others are exported by hand (1)
    • [ ] Most are exported by hand (0)
  3. Do you know how long data is kept, and can you see how it has changed?
    • [ ] Yes (2)
    • [ ] Partly (1)
    • [ ] No (0)
  4. Do you know which third-party services hold your data?
    • [ ] Yes, and we keep a list (2)
    • [ ] Roughly (1)
    • [ ] No (0)
  5. Could you move a key service to a different provider?
    • [ ] Yes (2)
    • [ ] With difficulty (1)
    • [ ] No (0)
  6. Can you trace a key figure back to its source data and the steps that produced it?
    • [ ] Yes (2)
    • [ ] For some figures (1)
    • [ ] No (0)
  7. Can you repeat a data collection (a survey, a cohort, a metric) the same way next time?
    • [ ] Yes (2)
    • [ ] Partly (1)
    • [ ] No (0)
  8. If the person who calculates a key number left tomorrow, could someone else produce it?
    • [ ] Yes (2)
    • [ ] With some effort (1)
    • [ ] No (0)
  9. Do different teams get the same answer from the same data?
    • [ ] Yes (2)
    • [ ] Usually (1)
    • [ ] No (0)
  10. Are data definitions kept close to the data, in the schema, metadata or a readme?
    • [ ] Yes (2)
    • [ ] Some of them (1)
    • [ ] No (0)
  11. Do access permissions reflect roles and regulatory requirements?
    • [ ] Yes (2)
    • [ ] Partly (1)
    • [ ] No (0)
  12. Would you notice if data went missing or was overwritten?
    • [ ] Yes, we have alerts (2)
    • [ ] Eventually (1)
    • [ ] No (0)
  13. Can people get the data they need without asking someone to pull it?
    • [ ] Yes (2)
    • [ ] Some of them (1)
    • [ ] No (0)
  14. Is a new metric easy to add?
    • [ ] Yes (2)
    • [ ] It takes a project (1)
    • [ ] No (0)

Score out of 28: 0-11, early; 12-21, developing; 22-28, established.

www.bayis.co.uk/checklists/data-maturity.html

Frequently asked questions

What is data maturity?

The degree to which an organisation can trust, find and use its data in consistent, repeatable ways. It runs from ad hoc storage and manual exports to data produced by processes, with known owners and written definitions.

Why does data maturity matter for AI?

A model works with the data it is given. At low maturity it inherits the same gaps and inconsistencies as the reports, and nobody can tell whether its output is right.

Do we need a data warehouse first?

Not necessarily. Many of the items here are practices: knowing where data lives, writing definitions down, keeping raw records. A single well-structured database handles tens of millions of rows without special effort.

How does this differ from the AI readiness assessment?

The assessment covers the whole AI decision: the goal, the data, its sensitivity, and who can judge the work. This checklist looks at the data in more depth.