
What is data sovereignty?
Data sovereignty is the principle that data is subject to the laws and jurisdiction of the nation where it is collected, stored, or processed — and, by extension, an organization's ability to control that legal exposure. Geography alone does not determine it.
Three things follow from that definition, and they are the whole subject in miniature:
Sovereignty is a question about jurisdiction, not about street addresses.
Jurisdiction attaches to the operator of the infrastructure as well as to the hardware itself.
Because of the first two, sovereignty is ultimately an evidence problem: you have to be able to prove where regulated data sits, whose law reaches it, and that it has not moved.
The second and third points are where the concept is usually lost. A definition that resolves sovereignty into a question of storage location is incomplete, and organizations that stop there are often confident about a compliance posture they cannot actually defend.
Data sovereignty, data residency, and data localization
These three terms are used interchangeably in sales decks, in procurement contracts, and occasionally in compliance documentation. They are not synonyms, and the conflation is the single most common source of sovereignty failure.
Dimension | Data residency | Data localization | Data sovereignty |
What it answers | Where does the data physically sit? | May the data leave the country? | Whose law can reach the data? |
What sets it | Contract, policy, or architecture choice | Statute or regulation | Jurisdiction of the data, the infrastructure, and the operator |
Who enforces it | Your vendor, under contract | A national regulator | Courts and law enforcement, potentially foreign ones |
Settled by geography alone? | Yes | Mostly | No |
Typical failure mode | Backups and replicas in an unapproved region | Cross-border transfer without required approval | Residency satisfied; provider still compellable abroad |
The compressed version, worth memorizing before your next vendor call: residency is where the data sits, localization is whether it may leave, and sovereignty is whose law can reach it.
For a deeper treatment of the three terms and the contractual language that blurs them, see our buyer's guide to data residency, sovereignty, and localization.
The Frankfurt test: Why residency ≠ sovereignty
The scenario. You buy a service from a US-headquartered provider. The provider runs a datacenter in Frankfurt and commits, in writing, that your data will reside there. Your data never leaves Germany.
The problem. The provider is incorporated in the United States, which means it may remain compellable under US legal process regardless of where the bytes physically sit. Residency achieved. Sovereignty, not necessarily.
The mechanism is the CLOUD Act, which requires a provider of electronic communication service or remote computing service to preserve, back up, or disclose customer data within its possession, custody, or control, regardless of whether that data is located within or outside the United States.¹ That phrase is worth memorizing, because it is the whole test: not where the data sits, but whether the provider can reach it. The Act does give providers a narrow route to move to quash where disclosure would conflict with the law of a qualifying foreign government, but that route is limited and does not extend to US persons.¹
The tension with European law is direct. Under the GDPR, answering a third-country authority's request is itself a transfer, and the European Data Protection Board is explicit that a request from a foreign authority does not in itself constitute a legal basis for processing or a ground for transfer.² A provider caught between the two is caught between two binding obligations.
The same logic applies in other directions. Jurisdiction can attach through the operator's incorporation, through its parent company, through the nationality of privileged administrators, and through subprocessors several layers down a chain nobody has mapped.
The useful diagnostic comes down to these three questions:
Where does the data sit? The residency question. Necessary, and the easiest of the three to answer.
Who operates the infrastructure? Not who owns the region label, but which legal entity holds administrative access and the encryption keys.
Whose law binds that operator? Including its parent, its subprocessors, and the jurisdictions its staff work from.
Sovereign cloud offerings are built to address these issues. The serious vendors combine a separate local legal entity, operator control held by nationals of the relevant jurisdiction, and customer-held encryption keys, so that a foreign demand cannot be satisfied even if it is made. The structures vary considerably in rigor, they are actively contested in legal commentary, and "sovereign" on a datasheet is not evidence that any of them are in place. Question 2 and question 3 are where diligence actually happens.
A note on terminology: Indigenous data sovereignty is a distinct concept, describing the right of Indigenous peoples to govern data about their communities, lands, and cultural knowledge, associated with the CARE Principles for Indigenous Data Governance.³ It is a separate and substantial body of work from the jurisdictional sense discussed here, and the two should not be conflated. A third usage, common in European dataspace initiatives, refers to a participant's ability to control the terms under which shared data is used.
What's driving sovereignty pressure
The pressure is not coming from one place, which is part of why it is hard to answer with a single policy.
Cross-border transfer rules. Under the GDPR, responding to a third-country authority's request is itself a transfer, subject to Chapter V in full. The standard the safeguards must meet is
protection essentially equivalent to that guaranteed inside the EEA, which turns transfer compliance from a paperwork exercise into a standing assessment of the receiving jurisdiction.²
Sector regimes. Health, financial services, and defense frequently carry requirements stricter than the general data protection law of the same country, and public-sector procurement usually sets the strictest bar of all.
National frameworks. China's Personal Information Protection Law and Cybersecurity Law together impose localization obligations on specified categories of data, with regulatory approval required before cross-border transfer. Requirements of this kind vary by country and by sector, so the rule that binds you is usually the sectoral one rather than the general data protection law.
The market response is now large enough to measure. Gartner forecasts worldwide sovereign cloud infrastructure-as-a-service spending reaching $80 billion in 2026, a 35.6% increase over 2025.⁴ Gartner estimates that geopatriation projects will shift roughly 20% of current workloads from global to local cloud providers.⁴ The geographic shape matters as much as the total: Europe's spending is forecast to rise from $6.9 billion in 2025 to $12.6 billion in 2026 and $23.1 billion in 2027, overtaking North America in 2027.⁴
The trajectory is steeper than the spend figures suggest. Gartner predicts that by 2030, more than 75% of European and Middle Eastern enterprises will geopatriate their virtual workloads into solutions designed to reduce geopolitical risk, up from less than 5% in 2025.⁵
Geopatriation (n.) — Gartner's term for moving company data and applications out of global public clouds and into local options such as sovereign clouds, regional providers, or an organization's own data centers, because of perceived geopolitical risk. Gartner named it one of its Top Strategic Technology Trends for 2026.⁵
However, none of those figures tell you whether the organizations doing the moving can demonstrate that it worked. That is the harder half of the problem, and it is where most programs are weakest.
What sovereignty actually requires of a governance team
Strip away the geopolitics and sovereignty resolves into five concrete requirements. None of them are policy questions. All of them are questions about whether your metadata is good enough to answer a regulator.
Know where every regulated dataset physically resides. Not by system, by dataset — including the copies. An inventory that covers your three warehouses but not the extracts feeding a regional reporting tool is not an inventory.
Know which jurisdiction attaches to it. Physical location plus the operator's jurisdiction plus the sectoral rule that applies to that data category. This is an attribute of the data, and it needs to live with the data.
Know whether lineage crosses a border you promised it wouldn't. Continuously, not at audit time. Pipelines change without anyone filing a ticket about jurisdiction.
Enforce geography-aware access controls. Who may query this dataset, from which jurisdiction, under what authority — applied to service accounts and AI agents as well as to people.
Produce audit evidence for all of the above. On demand, without a six-week fire drill.
The hard part is rarely writing the policy., since every modern organization has the policy. The hard part is proving the policy holds across a pipeline nobody fully mapped.
That distinction has legal teeth, not just rhetorical ones. The GDPR puts the burden on you rather than on the requesting authority: because Article 48 is not itself a ground for transfer, a controller facing a third-country order has to identify a lawful basis and a transfer ground elsewhere in Chapter V, and be able to show it.² Infrastructure gives you the capacity for compliance. Metadata gives you the evidence of it. Only one of those is an answer.
Requirement | Evidence a regulator will accept |
Regulated data inventory | A current, queryable catalog of datasets with sensitivity classification and owner, not a spreadsheet snapshot |
Jurisdictional attribution | Jurisdiction and applicable-regime tags applied at the dataset and column level |
No unapproved border crossings | End-to-end lineage showing every downstream destination, with the date the flow was last verified |
Geography-aware access | Access records tied to policy, showing who and what queried the data and under whose authority |
Continuous compliance | A time-stamped history of enforcement, not an annual attestation |
Each row maps to a capability rather than a process. A data inventory with jurisdictional tagging is a data catalog function. Identifying which of your thousands of datasets actually carry regulated exposure, so you are not attempting to govern all of them equally, is what Critical Data Manager is for. Detecting a flow that crosses a border you committed to is a lineage function, and no survey of engineering teams substitutes for it. And moving from periodic attestation to continuous enforcement is the premise of agentic data governance: declare the standard once, and let the system hold it as the estate changes. If your classification scheme is the weak link, start with what data classification actually requires.
Where residency promises quietly break
Residency commitments rarely fail at the primary store. They fail at the edges, in places configured for resilience or convenience by people who were not thinking about jurisdiction. Backups and disaster-recovery replicas are the ones worth checking first, because replication is configured for availability, which usually means geographic distribution.
The full list of places to look:
Backups and DR replicas. Configured for availability, not jurisdiction, and often the last thing anyone checks.
Archive and cold-storage tiers. Lifecycle policies move data to cheaper storage that may not carry the same regional guarantee as the hot tier.
CDN and edge caches. Content gets cached wherever it is requested from, by design.
Telemetry and product analytics. Usage data, logs, and traces flow to wherever the vendor's observability stack lives, and they frequently contain regulated fields nobody intended to send.
Vendor support access. A support engineer in a third country with production read access is a cross-border transfer, whether or not anyone characterized it that way.
Subprocessor chains. Your vendor's residency commitment is only as strong as the commitments it has obtained from its own vendors, several layers down.
Contract carve-outs for "product improvement" or model training. A residency clause with a training exception attached may be worth very little. Ask explicitly whether your data is used to train the vendor's models. The answer should be no, and it should be in the contract rather than in a blog post.
Six of the seven listed above are invisible from the console where you configured your region. They are visible in lineage, because lineage traces what the pipelines actually do rather than what the architecture diagram says they do. We go deeper on the detection problem in data sovereignty and cross-border sensitive data movement, and on designing for it upfront in data residency by design.
Metadata sovereignty: The exposure almost nobody audits
Here is the question that tends to change the conversation in a governance review.
You have carefully placed your regulated operational data in Frankfurt. Where are the classifications, the business definitions, the lineage graph, and the governance policies that describe that data?
For most enterprises, the honest answer is a vendor SaaS tenant whose jurisdiction was never evaluated, because metadata was not on anyone's list of regulated assets. But consider what that metadata actually contains: a structured map of every sensitive dataset you hold, what is in it, who can reach it, and how it flows. It is a detailed description of your regulatory exposure, and in some regimes it is arguably sensitive in its own right.
There are two distinct risks stacked here. The first is jurisdictional: metadata about regulated data, sitting under a legal regime you did not assess. The second is strategic, and it is the more expensive one. A vendor that holds your semantics holds your ability to migrate, to negotiate, and to adapt when the regulatory picture shifts again. You can sign a data processing agreement that pins your operational data to Frankfurt and still have traded a regulatory risk for a lock-in risk.
The principle we would argue for is a clean line between what you rent and what you own. Infrastructure is a procurement decision. Models are, increasingly, a procurement decision. You can rent infrastructure. You can rent models. You cannot afford to rent your enterprise knowledge. Your business definitions, classifications, lineage, and policy logic need to be architecturally yours, portable across clouds, jurisdictions, and tool stacks. That is what we mean by an independent knowledge layer, and it is a design principle rather than a feature.
Three questions to take into your next platform review:
In which jurisdiction does our governance metadata reside, and under whose operational control?
If we exited this platform in 18 months, what portion of our classifications, definitions, and lineage would leave with us in a usable form?
Is that metadata covered by the same residency and sovereignty commitments we negotiated for operational data, or was it never in scope?
Sovereignty in the age of AI agents
AI has not created a new category of sovereignty risk so much as it has multiplied the surface area and removed the paper trail. Three problems in particular:
1. The copy problem. Embeddings and vector stores are copies of your data. They are usually created outside the governed path, they frequently sit in a different service and sometimes a different region than the source tables, and they rarely appear in any inventory. A dataset pinned to an EU region whose embeddings live in a US-hosted vector database has been transferred, whatever the source-system configuration says. So what: your inventory has to cover derived representations, not just tables.
2. The crossing problem. Retrieval pipelines assemble context from multiple systems at query time, and inference may happen in a region chosen for model availability rather than for compliance. A single RAG request can cross two borders in under a second and leave behind a log line that records neither. So what: lineage has to extend through retrieval and inference, or the crossing is undetectable.
3. The authority problem. Agents act rather than answer. Most organizations can identify which agent touched a dataset; far fewer can name the human who authorized that agent to do it, or show that the authorization was still valid at the time. The regulatory clock here moved, and not in the direction most roadmaps assume: under the Digital Omnibus on AI, the EU AI Act's obligations for stand-alone high-risk systems now apply from 2 December 2027, and from 2 August 2028 for high-risk AI embedded in regulated products.⁶ That is a longer runway, not a lighter one, and the risk management, technical documentation, logging and human-oversight controls it asks for are exactly the capabilities that take eighteen months to build. So what: delegated authority needs to be recorded and expiring rather than implicit, and the recording needs to start well before the deadline. We've written about the chain of command for autonomous actions if you want the longer argument.
Registering every AI asset across every platform, and being able to produce that register on demand, is the practical starting point. That is the problem AI governance is built for.
Twelve questions to ask before you sign
Replace "are you compliant?" with these questions:
In which specific facilities will our data reside, and who owns and operates them?
Which legal entity holds our contract, and in which jurisdiction is it incorporated?
Does your parent company or any subprocessor introduce a jurisdiction we have not assessed?
Where do backups, DR replicas, and archive tiers live?
A good answer names a region per tier and offers to show the configuration. A vague answer here is the single strongest predictor of a failed audit.
Who holds the encryption keys, and can we hold them ourselves?
Can your employees access our data? Under what circumstances, from which countries, and with what audit trail?
A good answer includes break-glass procedures and a log we can review.
Is our data or metadata used to train your models, for product improvement, or for any AI feature?
The answer should be no, and it should be in the contract rather than in a policy page you can revise.
Where does our governance metadata reside, and is it covered by the same commitments as our operational data?
Can you support a fully isolated deployment with no cross-border flows, including for operational purposes?
Most vendors cannot. Far better to learn that now than during a market entry.
What certifications do you hold, and will you provide the audit reports rather than the summary? SOC 2 Type II, ISO 27001/27701, and where relevant HIPAA/HITECH and FedRAMP are baseline signals of operational maturity, not proof of compliance with your specific obligations.
What evidence can your platform produce for an auditor, and in what format?
What happens to our data and our metadata if we exit, and in what timeframe?
Sovereignty is a continuous discipline, not a certification
Regulations change. Your data footprint changes weekly. Your vendor relationships and their subprocessor chains change without notice. A point-in-time review conducted once a year cannot keep pace with any of that.
The organizations handling this well are not the ones with the thickest policy binder. They are the ones that treated sovereignty as an architectural discipline — designed in, instrumented, and continuously verified — and can therefore answer a regulator's question in an afternoon. That capability turns out to be a growth asset as much as a defensive one: it is what lets you enter a regulated market on schedule and win the enterprise deals that require a demonstrated posture rather than a promised one.
If you are evaluating whether your current catalog and governance layer can produce that evidence, let's talk.
Frequently asked questions
What is the difference between data sovereignty and data residency? Data residency is where data is physically stored — a geographic commitment. Data sovereignty is whose law can reach that data, which depends on the operator's jurisdiction as well as the hardware's location. You can satisfy residency in full and still lack sovereignty if your provider is compellable under a foreign government's legal process.
Can you have data sovereignty in a public cloud? Yes, with a provider incorporated and operated in the relevant jurisdiction. With a foreign-headquartered provider it depends on the specific legal and operational structure of its sovereign offering — separate legal entity, local operator control, customer-held encryption keys — and those structures vary considerably and remain actively contested.
Do backups have to meet the same data residency requirements? Yes, and this is one of the most commonly missed gaps. Backups, disaster-recovery replicas, and archive tiers frequently sit in different regions from primary storage, because replication is usually configured for resilience rather than jurisdiction. Verify each tier separately and get the answer in writing.
Which countries require data localization? Requirements vary by country and by sector, and are most common in financial services, healthcare, telecommunications, and the public sector. China's Personal Information Protection Law and Cybersecurity Law together impose localization obligations on specified data categories, with regulatory approval required before cross-border transfer. Many jurisdictions apply localization to specific data categories rather than universally, so check the sectoral rule that applies to you rather than the general data protection law.
Is Indigenous data sovereignty the same thing? No. Indigenous data sovereignty is a distinct concept: the right of Indigenous peoples to govern data about their communities, lands, and cultural knowledge, associated with the CARE principles. It is a separate body of work from the jurisdictional sense used in enterprise compliance, and the two should not be conflated.
Sources & notes
Every external claim on this page is independently verifiable. The public sources are listed here.
1. CLOUD Act § 103, codifying 18 U.S.C. § 2713 (required preservation and disclosure of data
within a provider's "possession, custody, or control, regardless of whether such
communication, record, or other information is located within or outside of the United
States") and amending 18 U.S.C. § 2703 to add comity analysis and motions to quash.
— U.S. Department of Justice, official text ↗
https://www.justice.gov/d9/pages/attachments/2019/04/09/cloud_act.pdf
2. A request from a third-country authority does not in itself constitute a legal basis for
processing or a ground for transfer (¶10); a disclosure in response to such a request is a
transfer under Chapter V (¶9); Article 48 is not a ground for transfer, and the controller
must identify one elsewhere in Chapter V (¶29); the standard for safeguards is protection
essentially equivalent to that guaranteed in the EEA (¶31 and fn.18, citing CJEU Case
C-311/18). — European Data Protection Board, Guidelines 02/2024 on Article 48 GDPR,
adopted 2 December 2024 ↗
https://www.edpb.europa.eu/system/files/2024-12/edpb_guidelines_202402_article48_en.pdf
3. The CARE Principles for Indigenous Data Governance (Collective Benefit, Authority to Control,
Responsibility, Ethics), published by the Global Indigenous Data Alliance in September 2019.
— Carroll et al., Data Science Journal 19(1):43, 2020 ↗
https://datascience.codata.org/articles/10.5334/dsj-2020-043
4. Worldwide sovereign cloud IaaS spending forecast at $80 billion in 2026, up 35.6% from 2025;
geopatriation projects estimated to shift 20% of current workloads from global to local cloud
providers; Europe forecast at $6,868M (2025), $12,587M (2026), $23,118M (2027), surpassing
North America in 2027 (Table 1). — Gartner press release, 9 February 2026 ↗
https://www.gartner.com/en/newsroom/press-releases/2026-02-09-gartner-says-worldwide-sovereign-cloud-iaas-spending-will-total-us-dollars-80-billion-in-2026
5. Gartner definition of geopatriation and its inclusion among the Top Strategic Technology
Trends for 2026; prediction that by 2030 more than 75% of European and Middle Eastern
enterprises will geopatriate their virtual workloads, up from less than 5% in 2025.
— Gartner press release, 20 October 2025 ↗
https://www.gartner.com/en/newsroom/press-releases/2025-10-20-gartner-identifies-the-top-strategic-technology-trends-for-2026
6. Date of application of Chapter III, Sections 1–3 set to 2 December 2027 for AI systems
classified as high-risk under Article 6(2) and Annex III, and to 2 August 2028 for those
classified under Article 6(1) and Annex I (recital 40). — Regulation (EU) 2026/1744 (Digital
Omnibus on AI), Official Journal, 24 July 2026 ↗
https://eur-lex.europa.eu/legal-content/EN/TXT/HTML/?uri=OJ%3AL_202601744
ANALYST ATTRIBUTIONS & DISCLAIMERS
Gartner, Forecast Analysis: Sovereign Cloud IaaS, Worldwide, as referenced in Gartner's press
release of 9 February 2026.
Gartner, Top Strategic Technology Trends for 2026 (Special Report), presented at Gartner IT
Symposium/Xpo, 20 October 2025.
Gartner does not endorse any vendor, product or service depicted in its research publications
and does not advise technology users to select only those vendors with the highest ratings or
other designation. Gartner research publications consist of the opinions of Gartner's research
organization and should not be construed as statements of fact. Gartner disclaims all
warranties, expressed or implied, with respect to this research, including any warranties of
merchantability or fitness for a particular purpose. GARTNER and Magic Quadrant are registered
trademarks and service marks of Gartner, Inc. and/or its affiliates in the U.S. and
internationally and are used herein with permission. All rights reserved.
- AI
- Cloud Transformation
- Data Catalog
- Data Quality
- Digital Transformation
- Enterprise Data Catalog
- Modern Data Stack
Keep reading
More from the data desk



