Sovereign AI for UK organisations
Many organisations cannot send their data to a third-party model API. The data is personal, or it belongs to clients, or a regulator restricts where it goes, or a leak would cost more trust than the organisation can afford to lose. Open-weight models run wherever the data already is, which takes the model provider out of the arrangement.
This page sets out what sovereignty means in practice, who needs it, and the three ways we deploy marigold (our open-source inference platform) to provide it. (To see where your organisation stands first, take the AI readiness assessment; to go straight to a conversation, discuss a deployment.)
What does "sovereign AI" mean for an organisation?
Sovereign AI means an organisation controls four things: where its data is processed, who can see it, which model runs, and which laws apply. Each has a precise answer, and the answers change with the deployment.
Where the data is processed. Residency is the physical location of the servers that store and process the data. For UK residency, that means hardware in the UK: your own premises, or a UK cloud region such as AWS London (eu-west-2).
Who can see it. Sending a prompt to a model API gives the provider a copy of the data, whatever its policy says about retention and training. Running an open-weight model removes that provider; whoever operates the hardware remains.
Which model runs. Hosted APIs update and retire models on the provider's schedule. An open-weight model is a set of files you hold, and the version you test stays the version you run until you decide to change it.
Which laws apply. Data stored in the UK falls under UK law, and residency alone does not remove other jurisdictions. AWS, Microsoft and Google are US companies, and the US CLOUD Act allows US authorities to require a US provider to produce data in its possession, custody or control, wherever that data is stored. On hardware you own, no cloud provider is involved, so there is no third party for a foreign authority to compel.
Who needs sovereign AI?
Any organisation whose data cannot go to a third-party model provider, whether the reason is the law, a contract, or the trust of the people the data describes.
Should we host our own inference infrastructure?
Yes, if the data cannot leave your boundary; otherwise the answer depends on volume. A hosted API is the sensible default for data with no sensitivity constraint and irregular usage. Private inference (running open-weight models on infrastructure you control) pays off when the data cannot go to a third party at all, or when usage is steady enough that dedicated hardware costs less than per-token pricing, which usually happens above a few million tokens a day.
The engineering questions that follow (which GPUs, which serving software, how to scale) are covered in AI inference infrastructure: a decision framework, and the running costs in what private inference costs.
What is marigold?
marigold is our open-source inference platform for open-weight models such as Llama, Mistral and Qwen. It offers the same API as OpenAI and Anthropic, so software written for those services can point at marigold without changes, and it runs the same way on a single server, in a cloud account, or as a hosted service.
Its models cover text generation, embeddings, image understanding and generation, text-to-speech, and evaluation. (For engineers: marigold.run, the source on GitHub, the tutorials, and how marigold works.)
Which deployment option fits?
The software is the same in all three options. They differ in whose hardware it runs on, and so in how far sovereignty extends. Moving between them changes the infrastructure and leaves your applications as they are.
On your own hardware
marigold runs on servers you own, on your premises. It needs an internet connection only to download the models; after that it runs with no connection at all. No cloud provider is involved, so no third party holds the data. This is the strongest option.
In your cloud account
We deploy marigold into your own cloud account, inside the security setup you already run: your network rules, access controls, logging and billing. Costs and data egress appear in the monitoring you already have. Your cloud provider's jurisdiction applies.
marigold needs only compute instances and storage, so it runs on any of the major clouds. We deploy to AWS with Terraform today; adding Azure or Google Cloud takes about a week, on request.
A deployment runs in four stages:
- A readiness check of the data, its sensitivity and the use case.
- Deployment and configuration in your account.
- Integration with one agreed system or workflow.
- Handover, with documentation.
Hosted by us
We run marigold on AWS in the region you choose and bill you directly. You receive an API endpoint and manage no infrastructure, and model versions change only by agreement with you. AWS is a US company, so the CLOUD Act point above applies.
| Your hardware | Your cloud account | Hosted by us | |
|---|---|---|---|
| Where data is processed | Your premises | The region you choose | The AWS region you choose |
| Model providers see the data | No | No | No |
| Who operates the infrastructure | You | You, on your cloud provider | Bay Information Systems, on AWS |
| Model versions | Pinned by you | Pinned by you | Pinned by agreement with you |
| Outside US jurisdiction | Yes | Only with a non-US provider | No |
| Internet connection | Only to download models | Required | Required |
What does getting it wrong cost?
Four problems come up repeatedly. A data incident involving personal or client data counts as your breach even when a processor caused it (the ICO's definition covers data sent to a third party). A provider can change or retire a model that a process depends on, on its own schedule. A single vendor's API sets your prices, and leaving means rewriting against another. And a regulator or client can ask where their data went, at which point the answer needs to be specific. (The readiness assessment includes a question on which of these apply to your data.)
Frequently asked questions
Are open-weight models good enough?
For most document, search and extraction tasks, yes. Smaller open-weight models handle classification, extraction, summarisation and question-answering over documents with results comparable to much larger models, and the largest open-weight models come close to frontier quality. The leading closed models usually keep an edge on the hardest open-ended reasoning and on long, multi-step agent tasks.
Is it more expensive than using an API?
It depends on volume. At low or irregular volumes a hosted API usually costs less; at steady volumes above a few million tokens a day, dedicated hardware usually costs less. The break-even is worth calculating for your own workload. (What private inference costs.)
Can we still use ChatGPT or Copilot for some things?
Yes. Decide by data sensitivity: hosted tools for data with no constraint, private inference for the rest. A short written policy on what staff may paste into hosted tools covers most of the risk. (AI readiness assessment.)
Where exactly is our data stored and processed?
On your premises, for the on-premises option. For the other two options, in the cloud region chosen at deployment; for UK residency, that is a UK region such as AWS London (eu-west-2).
AWS is a US company. Does that matter?
It can. Data in AWS London stays in the UK, and US law (the CLOUD Act) can still require AWS to produce data it holds. If that matters for your data, run marigold on your own hardware, or in an account with a cloud provider outside US jurisdiction.
Discuss a deployment
Tell us what the data is and which option interests you, and we will reply within one working day. (What we collect is set out in the privacy notice.)