Do you need private inference?

Private inference means running AI models on infrastructure you control, so the data never goes to a model provider. Whether you need it comes down to three risks (regulation, exposure, and the reuse of submitted data) and two practical questions: which jurisdictions can reach the data, and whether the system has to run offline. This checklist works through them and suggests which kind of deployment fits.

How does the checklist work?

Eight questions, about five minutes. Each answer shows guidance and a link to further reading. Seven questions are scored: the higher the score, the more freedom you have to use a hosted model API. A regulatory or contractual yes costs more points than the other answers, because either one usually rules out a hosted API on its own. The eighth question, on usage, affects only the cost case.

The questions

1. Does a law, regulator or framework restrict where any of the data can be processed?

Yes (for example NHS data frameworks, FCA rules, legal professional privilege)

Then the question is how to run AI compliantly. A hosted model API is a third-party processor, and the answer for regulated data is usually no.

Not sure

Find out before choosing a tool. Your data protection officer or compliance lead can usually answer quickly, because the frameworks are specific about third-party processing.

No

Then regulation does not decide this, and the questions below do.

2. Does the data include personal data about customers, members, patients or staff?

Yes

UK GDPR applies to every processor the data passes through, a model provider included, and processing outside the UK brings the international transfer rules into play.

Some of it

Consider whether it can be found and removed before processing. Detection misses some cases, so redaction reduces the risk without removing it.

No

One constraint fewer. The remaining questions still apply.

3. Do client contracts or confidentiality agreements restrict third-party processing?

Yes

Sending the data to a model provider may breach them, whatever the provider's own terms say. Read the clauses before choosing a tool.

Not sure

Check for clauses on subcontractors, processors and data location. They often cover AI services without naming them.

No

Then contracts do not decide this.

4. If this data appeared somewhere it should not, what would it cost you?

A great deal: client trust, legal action or regulatory attention

Then exposure is the deciding risk. A breach at your processor is a breach you may have to report.

Embarrassing, but survivable

Weigh it against the convenience of a hosted service. Keeping only the most sensitive subsets in-house is often enough.

Little or nothing

A hosted API is a reasonable choice for this data.

5. Do you know whether your AI services can use submitted data for training?

Yes, and our terms exclude it

Keep a record of the terms you relied on; providers revise them.

Staff use consumer chat tools, and we have not checked

Consumer chat services may use conversations for training unless a setting or plan excludes it. Check the settings, and set a policy on what staff may paste in.

We do not use an AI service yet

Then you can choose on these criteria from the start. Running an open-weight model involves no model provider at all.

6. Would it matter if a foreign government could require your cloud provider to hand over the data?

Yes

Then UK residency is not enough on its own. US law can reach US cloud providers wherever the data is stored; your own hardware, or a provider outside US jurisdiction, avoids that.

Not sure

It usually matters for defence, security and some public-sector work, and for clients who ask. Check what your clients' contracts say about data location and access.

No

Then a UK region with any major cloud provider meets the residency requirement.

7. Does the system need to run without an internet connection?

Yes

Then it has to run on your own hardware. marigold runs with no connection once its models are downloaded.

No

Any of the deployment options will work.

8. How steady is the expected usage?

Steady and high (millions of tokens a day)

At this volume, dedicated hardware may cost less than per-token pricing, whatever the sensitivity. The break-even is worth calculating.

Occasional or unpredictable

Per-token pricing suits irregular use, where dedicated hardware would sit idle. If private inference is needed for other reasons, a hosted private service spreads the hardware cost.

We do not know yet

Estimate it from the number of documents or requests a day; the cost comparison depends on it.

Your results

Scoring runs in your browser and needs JavaScript. The printable checklist below lists every question and answer.

Can I print the private inference checklist?

Yes. The version below lists every question with the points for each answer. Questions marked "no points" change the advice without changing the score.

  1. Does a law, regulator or framework restrict where any of the data can be processed?
    • [ ] Yes (for example NHS data frameworks, FCA rules, legal professional privilege) (0)
    • [ ] Not sure (3)
    • [ ] No (4)
  2. Does the data include personal data about customers, members, patients or staff?
    • [ ] Yes (0)
    • [ ] Some of it (1)
    • [ ] No (2)
  3. Do client contracts or confidentiality agreements restrict third-party processing?
    • [ ] Yes (0)
    • [ ] Not sure (3)
    • [ ] No (4)
  4. If this data appeared somewhere it should not, what would it cost you?
    • [ ] A great deal: client trust, legal action or regulatory attention (0)
    • [ ] Embarrassing, but survivable (1)
    • [ ] Little or nothing (2)
  5. Do you know whether your AI services can use submitted data for training?
    • [ ] Yes, and our terms exclude it (2)
    • [ ] Staff use consumer chat tools, and we have not checked (0)
    • [ ] We do not use an AI service yet (2)
  6. Would it matter if a foreign government could require your cloud provider to hand over the data?
    • [ ] Yes (0)
    • [ ] Not sure (1)
    • [ ] No (2)
  7. Does the system need to run without an internet connection?
    • [ ] Yes (0)
    • [ ] No (2)
  8. How steady is the expected usage? (no points)
    • [ ] Steady and high (millions of tokens a day)
    • [ ] Occasional or unpredictable
    • [ ] We do not know yet

Score out of 18: 0-8, private inference from the start; 9-14, private inference in your cloud; 15-18, a hosted api is likely fine.

www.bayis.co.uk/checklists/private-inference.html

Frequently asked questions

What is private inference?

Running AI models on infrastructure you control (your own hardware, your cloud account, or a dedicated hosted service), so prompts and data never reach a model provider such as OpenAI, Anthropic or Google.

Is UK data residency enough?

It places the data under UK law. It does not stop US law from reaching a US cloud provider, which matters for some data and not for other data.

Do AI providers train on the data we send them?

It depends on the service and its terms. Business API and enterprise terms usually exclude training by default; consumer chat services and some free tiers may use submitted data unless a setting or plan excludes it. Terms change, so keep a record of the ones you relied on.

Are open-weight models good enough to replace a hosted API?

For most document, search and extraction tasks, yes. The leading closed models usually keep an edge on the hardest open-ended reasoning and on long, multi-step agent tasks.