← Back to Articles

Where Does Your Data Actually Go? A Sensitivity Map for Choosing Local, EU Cloud or Public AI

Picture an employee who knows AI well, in a company that has decided Copilot is the official tool. If they think another tool, Claude for example, does a better job for what they need, there’s a good chance they’ll use it anyway, on the side, without telling anyone. That’s what shadow AI looks like in a lot of companies, and it isn’t necessarily the least-skilled people doing it. In this case it’s the one who knows AI best.

Some companies react by forbidding AI altogether. I don’t think anyone can afford that. A company that drops AI tools immediately loses a huge competitive advantage: in intelligence and thinking capacity, in pace of work, in access to expertise that employees, and even C-level executives, don’t really have otherwise. The most efficient companies of the next wave are flexible, use agents widely in every department and keep updating their ways of working with the most advanced capabilities AI can give them. Avoiding AI today is not an option. It’s use it or die, in the coming months, or years if you’re lucky.

So the question to settle is which data can go into which tool. What works in practice is a map simple enough that everyone can apply it in a few seconds, and the four steps below build one.

Step 1: Sort your data into four tiers

You don’t need a classification project. Four buckets are enough, as long as each one comes with examples from your own business.

  1. Public. Anything already on your website, in brochures or in public sources: marketing copy, published prices, general industry research.
  2. Internal. Everyday information that would be awkward to leak without really hurting you: routine meeting notes, process documents, drafts of ordinary emails.
  3. Confidential. What gives you an edge or what you’ve promised to protect: client files, contracts, financial figures, R&D, pricing models, source code, strategy documents, product plans.
  4. Personal and special-category. Anything about identifiable people, and above all the special categories the GDPR protects most strictly, such as health data. Treat other sensitive personal information the same way: financial situations, HR cases and anything a client told you in confidence.

A document that mixes tiers takes the highest one.

In practice, the clients I talk to don’t make these distinctions. For them it’s just “our data”. When they do get protective, it’s about two things: financial figures and R&D. That’s a good place to start the conversation, because everyone already agrees those two belong in tier 3.

I’d add one rule that sits outside the tiers: keys, passwords and any kind of access credentials never go into an AI tool, whatever the tier and whatever the deployment. It’s the one strict rule I apply on my own sites.

Step 2: Know the four deployment options

Most writing about local AI presents two choices, the cloud or your own server. There are four, and the two in the middle are where most SMBs end up.

  1. Consumer AI tools. Free or personal-plan chatbots. Their terms often let the provider use your inputs to improve its models unless you switch that off, and you have almost no contractual protection.
  2. Business-tier cloud AI. Team or enterprise plans from the large providers, with a contractual commitment not to train on your data, a data processing agreement and admin controls. Processing may still happen outside the EU.
  3. EU-hosted or sovereign cloud AI. Services that keep processing inside the EU, run by a European entity or under an EU data boundary. This removes most transfer questions and, when the provider isn’t subject to US jurisdiction, the CLOUD Act exposure as well.
  4. Local or on-premise AI. Open-weight models on hardware you control. The data stays on your network, and so does the full responsibility for security, updates and uptime.

Step 3: Match tiers to options

This is the map. Adapt it to your company and put it in your acceptable-use policy.

Data tier Consumer AI Business-tier cloud EU-hosted cloud Local
Public Yes Yes Yes Yes
Internal No Yes Yes Yes
Confidential No Case by case, with contract review Yes Yes
Personal / special-category No Only with DPA, legal basis and transfer review Yes, with DPA and legal basis Yes, with legal basis and proper security

Consumer tools are acceptable only for public data, and that one rule removes most of the shadow-AI risk on its own. Look at the local column as well: it’s never an automatic yes, because running on your own hardware gives you neither a legal basis to process personal data nor a secured machine.

Step 4: What regulation does and doesn’t change

This is where most local-AI marketing gets careless, so it’s worth being precise.

The GDPR applies wherever the model runs. Local processing avoids international transfers and cuts the number of processors you depend on. Your obligations stay as they were: a legal basis, purpose limitation, data minimisation, security and the rights of the people whose data you hold.

Erasure is hard everywhere. Once a model has been fine-tuned on personal data, removing one person from it is technically difficult, whether the model sits in your office or in a data centre. The European Data Protection Board’s Opinion 28/2024 on AI models says a model trained on personal data can’t simply be assumed to be anonymous. The practical consequence is to keep personal data out of training and fine-tuning altogether and bring it in at query time, by letting the model search your documents, so that deleting a file really does delete it.

The EU AI Act looks at use, not location. Your obligations depend on what you use AI for and on whether you’re a provider or a deployer. One of them already applies to nearly every company using AI: since February 2025, deployers have to ensure sufficient AI literacy among the staff who use it.

Vendor risk is where location matters most. Terms of service and prices change, and providers get acquired. A provider under US law can be compelled to hand over data under the CLOUD Act even when it’s stored in Europe, and that’s the strongest honest argument for EU-hosted and local options. The NIST AI Risk Management Framework and ISACA’s AI audit guidance give you a structured way to assess it, which beats relying on a vendor’s sales deck.

Put it in place this month

Roles, approval workflows and the audit trail are covered in the Governance and Compliance series in my Playbook. If this is part of a wider rollout, it belongs in the first 30 days of The AI Adoption Sprint.

Of the five, the third item usually does the most work: once people have an approved tool for confidential data, most of them stop reaching for the free one.