The question is usually framed wrongly. “Public AI or your own model?” sounds like a matter of principle – cloud or sovereignty, convenience or control. In practice it is not an either-or question but a matter of assignment: which task belongs on which model, and who decides that by what rule? Sovereign AI in sector7’s sense is therefore not cloud purism. It is the discipline of placing every model where it belongs.
The reflex to run everything locally for data protection reasons is just as expensive as the opposite reflex of handing everything to the public cloud out of convenience. Both confuse an architectural decision with a worldview. Anyone who wants to work sovereignly makes the decision per task – and makes it verifiable.
Plan in the public cloud, execute locally
A usable rule of thumb separates by what happens to the data. Public cloud models are superior wherever the task is research, drafts, ideation, and broad world knowledge – tasks where no content worth protecting goes into the prompt. Call it “planning on the internet”: you have a type of contract explained, sketch out an argumentation structure, or weigh options against one another. The input is uncritical, and the performance of the large proprietary models is real here.
Execution on sensitive content belongs in a different place. As soon as a model accesses your contracts, personnel data, design documents, or customer files, the calculus changes. This content should not leave your premises – here a self-hosted, open model works on your own or sovereignly operated infrastructure in Germany. “Execute locally” means: the actual work on the data worth protecting takes place where you retain control over storage location, access, and jurisdiction.
Why this now holds up in practice
Two years ago, this split was a compromise at the expense of quality. Today it rarely is. The gap between open and proprietary models has narrowed considerably on many production-relevant tasks – coding, knowledge questions, summarization. Open models such as DeepSeek V4, Qwen 3.5, Llama 4, or Mistral Medium 3.5 achieve scores on common benchmarks that were, until recently, reserved for the proprietary top tier: DeepSeek V4 Pro, for instance, scores 80.6 on SWE-Bench Verified and 90.1 on GPQA Diamond with a context window of one million tokens.
Soberly, it must be noted where the gap persists: on demanding, multi-step reasoning, on long agentic work chains, and on reliability in edge cases, the proprietary top models still lead – the lead is shrinking but has not disappeared. This is precisely why the mix is not a stopgap but the factually correct answer: you use the public top tier where it counts, and the local model where the data situation demands it and the open quality suffices.
Residency is not sovereignty
A widespread misconception deserves a clear correction here, because it leads to expensive wrong decisions. Running a model in an AWS or Azure Frankfurt region gives you data residency – the data sits geographically in Germany. That is not sovereignty. Residency answers the question “Where does the data sit?”, sovereignty the question “Whose law can reach it, and who can compel disclosure?”.
The difference is legally concrete. The US CLOUD Act obligates US companies to hand over data on the order of US authorities – regardless of whether the servers stand in Frankfurt or Virginia. A US provider with a German data center is subject to both legal orders. The Data Act, applicable EU-wide since September 12, 2025, conversely requires cloud providers to take technical precautions against unlawful third-country access to data stored in the EU. The two sets of rules impose opposing obligations on the same provider. Anyone who needs genuine sovereignty cannot avoid an operator subject exclusively to EU law – not merely a data center location. For sensitive execution, this is the reason to run the local model on European-governed infrastructure.
The gateway turns model choice into a decision
Between “plan in public” and “execute locally” stands the genuinely difficult task: the steering. Without a common control point, chance decides in practice which model receives a request – the developer who happens to have the most convenient API key at hand, or the tool with the preconfigured default. In this way sensitive content ends up at the wrong model, not out of malice but for lack of a rule.
A governed AI gateway solves this by bundling all model traffic through one access point and routing it according to fixed rules. The criteria can be named:
- Sensitivity: content classified as confidential goes to the local model, uncritical tasks to the most capable suitable public one.
- Cost and token budget: short classification and routine tasks run on smaller, cheaper models; complex reasoning is deliberately routed to a strong model. Publicly available router methods show that a substantial part of the cost can be saved this way without appreciably lowering the quality of results.
- Traceability: every request is logged, recurring requests are cached, budgets are enforced. The rule is defined once and applies to every connected application.
The decisive effect is governance-driven, not technology-driven. Routing is not a configuration knob but a governed surface: model choice becomes a documented decision instead of a by-product. And because all applications speak through the same access point, switching a provider or model becomes a configuration change – not a rebuild of every single integration. This is exactly the most effective protection against lock-in: not forgoing a particular provider, but the freedom to swap it out at any time.
What is genuine value here – and what is marketing
Two points need to be kept apart. Sovereign AI is occasionally sold as forgoing performance – “local, but more modest.” That was once true and hardly is anymore: the open middle tier is enough for a large part of the execution on your own data. Conversely, “sovereign cloud” is often stuck as a label on offerings that deliver residency but not sovereignty. Both simplifications are misleading.
The dependable core is unspectacular: sort your tasks by sensitivity, choose the suitable model per class, and place a control point in front that enforces and logs this assignment. This is less an AI question than a question of clean architecture and lived governance – the same discipline with which one also segments networks and controls access.
How sector7 supports you
We build exactly this split as an operated setup, not a slide deck. Our offering Sovereign AI begins with governance – use-case and risk assessment per the EU AI Act, data hygiene, connection to existing ISO 27001 and NIS-2 structures – and leads to secured operation: secured private language models and managed RAG on your documents, run on self-hosted open models in our own server park in Germany. The steering in between we set up as a gateway with rule-based routing: sensitive content local, uncritical content to the suitable public model – with cost, token, and logging transparency.
Deliberately not part of this is the training of our own foundation models; we rely on proven open and public models and bring them safely into use. For very large compute capacity we bring in sovereign partners. On request we take over ongoing operation from our data center – binding SLA models for this we are currently building and align the scope and commitments concretely with you case by case. At the start there is not a model but a conversation: we sort your plans with an open mind and implement what brings value.
This article is a professional assessment and does not replace legal advice in individual cases.
Sources
- https://www.jamesm.blog/ai/state-of-open-weight-models-2026/
- https://whatllm.org/blog/open-source-vs-proprietary-llms-2025
- https://www.vellum.ai/open-llm-leaderboard
- https://blogs.vmware.com/cloud-foundation/2025/11/18/the-great-cloud-charade-why-data-residency-isnt-data-sovereignty/
- https://particula.tech/blog/eu-ai-act-data-sovereignty-residency
- https://beyondscale.tech/blog/ai-data-residency-sovereignty-gdpr-cloud-act
- https://www.digitalapplied.com/blog/llm-model-routing-2026-cost-quality-optimization-engineering-guide
- https://neuraltrust.ai/blog/llm-model-routing