Sovereign AI for UK organisations

Many organisations cannot send their data to a third-party model API. The data is personal, or it belongs to clients, or a regulator restricts where it goes, or a leak would cost more trust than the organisation can afford to lose. Open-weight models run wherever the data already is, which takes the model provider out of the arrangement.

This page sets out what sovereignty means in practice, who needs it, and the three ways we deploy marigold (our open-source inference platform) to provide it. (To see where your organisation stands first, take the AI readiness assessment; to go straight to a conversation, discuss a deployment.)

What does "sovereign AI" mean for an organisation?

Sovereign AI means an organisation controls four things: where its data is processed, who can see it, which model runs, and which laws apply. Each has a precise answer, and the answers change with the deployment.

Where the data is processed. Residency is the physical location of the servers that store and process the data. For UK residency, that means hardware in the UK: your own premises, or a UK cloud region such as AWS London (eu-west-2).

Who can see it. Sending a prompt to a model API gives the provider a copy of the data, whatever its policy says about retention and training. Running an open-weight model removes that provider; whoever operates the hardware remains.

Which model runs. Hosted APIs update and retire models on the provider's schedule. An open-weight model is a set of files you hold, and the version you test stays the version you run until you decide to change it.

Which laws apply. Data stored in the UK falls under UK law, and residency alone does not remove other jurisdictions. AWS, Microsoft and Google are US companies, and the US CLOUD Act allows US authorities to require a US provider to produce data in its possession, custody or control, wherever that data is stored. On hardware you own, no cloud provider is involved, so there is no third party for a foreign authority to compel.

Who needs sovereign AI?

Any organisation whose data cannot go to a third-party model provider, whether the reason is the law, a contract, or the trust of the people the data describes.

Membership networks and charities Personal data about members and donors, held by organisations whose reputation rests on discretion. Finding personal data in text
Health Health data is special category data under UK GDPR, and NHS data frameworks restrict where it can be processed. Why private inference
Research and professional services Client-confidential archives. Searching them with AI through a hosted API sends the documents to a third party at indexing and again at query time. RAG risks and costs
Insurance and finance Claims, policies and customer records under FCA conduct rules and client confidentiality. Running AI inside your own infrastructure
Defence and security suppliers Data that cannot leave a controlled environment. marigold runs with no network connection once its models are installed. Deployment tiers for open-weight models

Should we host our own inference infrastructure?

Yes, if the data cannot leave your boundary; otherwise the answer depends on volume. A hosted API is the sensible default for data with no sensitivity constraint and irregular usage. Private inference (running open-weight models on infrastructure you control) pays off when the data cannot go to a third party at all, or when usage is steady enough that dedicated hardware costs less than per-token pricing, which usually happens above a few million tokens a day.

The engineering questions that follow (which GPUs, which serving software, how to scale) are covered in AI inference infrastructure: a decision framework, and the running costs in what private inference costs.

What is marigold?

marigold is our open-source inference platform for open-weight models such as Llama, Mistral and Qwen. It offers the same API as OpenAI and Anthropic, so software written for those services can point at marigold without changes, and it runs the same way on a single server, in a cloud account, or as a hosted service.

Its models cover text generation, embeddings, image understanding and generation, text-to-speech, and evaluation. (For engineers: marigold.run, the source on GitHub, the tutorials, and how marigold works.)

Which deployment option fits?

The software is the same in all three options. They differ in whose hardware it runs on, and so in how far sovereignty extends. Moving between them changes the infrastructure and leaves your applications as they are.

On your own hardware

marigold runs on servers you own, on your premises. It needs an internet connection only to download the models; after that it runs with no connection at all. No cloud provider is involved, so no third party holds the data. This is the strongest option.

Who runs it: your team. Price: the software is open source and free; help with hardware sizing, installation and configuration is priced per engagement. (Sizing guidance: GPU options for self-hosted inference.)

In your cloud account

We deploy marigold into your own cloud account, inside the security setup you already run: your network rules, access controls, logging and billing. Costs and data egress appear in the monitoring you already have. Your cloud provider's jurisdiction applies.

marigold needs only compute instances and storage, so it runs on any of the major clouds. We deploy to AWS with Terraform today; adding Azure or Google Cloud takes about a week, on request.

A deployment runs in four stages:

  1. A readiness check of the data, its sensitivity and the use case.
  2. Deployment and configuration in your account.
  3. Integration with one agreed system or workflow.
  4. Handover, with documentation.

Who runs it: you, after handover. Price: a fixed fee for the readiness check, then a fixed price for the deployment, quoted from it.

Hosted by us

We run marigold on AWS in the region you choose and bill you directly. You receive an API endpoint and manage no infrastructure, and model versions change only by agreement with you. AWS is a US company, so the CLOUD Act point above applies.

Who runs it: Bay Information Systems. Price: monthly, based on usage, and available on request. (Include your expected volume when you get in touch.)

The three options against the four questions.
Your hardware Your cloud account Hosted by us
Where data is processed Your premises The region you choose The AWS region you choose
Model providers see the data No No No
Who operates the infrastructure You You, on your cloud provider Bay Information Systems, on AWS
Model versions Pinned by you Pinned by you Pinned by agreement with you
Outside US jurisdiction Yes Only with a non-US provider No
Internet connection Only to download models Required Required

What does getting it wrong cost?

Four problems come up repeatedly. A data incident involving personal or client data counts as your breach even when a processor caused it (the ICO's definition covers data sent to a third party). A provider can change or retire a model that a process depends on, on its own schedule. A single vendor's API sets your prices, and leaving means rewriting against another. And a regulator or client can ask where their data went, at which point the answer needs to be specific. (The readiness assessment includes a question on which of these apply to your data.)

Frequently asked questions

Are open-weight models good enough?

For most document, search and extraction tasks, yes. Smaller open-weight models handle classification, extraction, summarisation and question-answering over documents with results comparable to much larger models, and the largest open-weight models come close to frontier quality. The leading closed models usually keep an edge on the hardest open-ended reasoning and on long, multi-step agent tasks.

Is it more expensive than using an API?

It depends on volume. At low or irregular volumes a hosted API usually costs less; at steady volumes above a few million tokens a day, dedicated hardware usually costs less. The break-even is worth calculating for your own workload. (What private inference costs.)

Can we still use ChatGPT or Copilot for some things?

Yes. Decide by data sensitivity: hosted tools for data with no constraint, private inference for the rest. A short written policy on what staff may paste into hosted tools covers most of the risk. (AI readiness assessment.)

Where exactly is our data stored and processed?

On your premises, for the on-premises option. For the other two options, in the cloud region chosen at deployment; for UK residency, that is a UK region such as AWS London (eu-west-2).

AWS is a US company. Does that matter?

It can. Data in AWS London stays in the UK, and US law (the CLOUD Act) can still require AWS to produce data it holds. If that matters for your data, run marigold on your own hardware, or in an account with a cloud provider outside US jurisdiction.

Discuss a deployment

Tell us what the data is and which option interests you, and we will reply within one working day. (What we collect is set out in the privacy notice.)