Yes, you can use AI on your own data, but not on all of it. Three questions settle it: which data is involved, where is it processed, and can you account for what happened afterwards. The classification is the work. The rest is contracts and settings.
The question always arrives in the same shape. Somebody wants to use AI on real work, and somebody else asks whether that is actually allowed.
Then it stays unresolved, and meanwhile five employees use a free account for the same thing without anyone having decided. That is the situation that is genuinely risky, not the controlled use.
This guide is the short route to a decision.
The three questions that settle it
Which data? Personal data is a different conversation from business data. The line falls on whether something is about a person or about the company.
Where is it processed? That is a question of geography and of who can be compelled to disclose, not of the vendor's nationality.
Can you account for it afterwards? In practice both GDPR and the EU AI Act are mostly about this. Not about having done it perfectly, but about being able to show what you did and why.
Answer those three clearly and you are ahead of most.
Step 1: Classify before you do anything else
There are four levels, and they determine everything that follows:
- PUBLIC can go anywhere.
- INTERNAL to systems you control.
- CONFIDENTIAL only with written conditions on place of processing, no training, and an audit trail.
- RESTRICTED is never sent.
The rule of thumb on the hard boundary turns on two things: personal data is always RESTRICTED, and so are exact figures such as revenue, margin and EBITDA. Descriptions, terms and assessments are CONFIDENTIAL.
The exercise itself takes half an hour for most companies, and examples of what belongs where are in the four classification levels.
What matters is that the classification sits on the source. If it sits on the individual request, somebody has to judge every time, and that judgement gets made wrong on a busy Friday.
Step 2: Establish where processing happens
Three questions for the vendor, whether or not you pay for the product:
- Which country holds the data centre that processes our requests?
- Is our data used to train models? The answer belongs in the contract, not in an email.
- Who are the subprocessors, and where are they?
Note that "we are GDPR compliant" answers none of the three. It is a claim about an outcome, not information about a place.
A US provider processing in an EU data centre can be perfectly fine. A European provider forwarding to a US subprocessor is not. Why place rather than nationality is what counts is developed in AI and data security in the EU.
Step 3: Make sure something enforces the rule
A classification that exists only in a document does not get followed. Something has to sit between your data and the model.
In practice there are three layers: fields can be stripped categorically based on which kind of model is at work, a scanner can catch keys and passwords in free text, and personal data can be handled on the way to the model.
Only the first layer is categorical. The other two are pattern matching and can err in both directions. Which is why classification is the responsibility and the other layers are the safety net.
Ask specifically: which fields do you never send, and is that a setting or is it built in? If it is a setting, it is not really a guarantee.
Step 4: Keep enough to account for it
This is the part people skip, and it is the part that gets asked about.
Four lines per analysis is enough: which model, which data in, what came out, and what you used versus discarded. One folder per strategy cycle.
For most companies using AI for analysis rather than for decisions about people, that is substantially the whole obligation under the EU AI Act. What applies beyond it, and when you are at the heavy end, is covered in the EU AI Act for SMEs.
How it plays with GDPR day to day, including what gets stopped on the way in and out, is in GDPR and AI-assisted strategy.
The four mistakes that cost most
Assuming EU hosting is enough. It settles where, not what. An unanonymised customer list is still personal data in an EU data centre.
Assuming anonymisation is enough. If the recipient can identify the person from context, it is still personal data. "Customer segment A" is anonymous. "Our largest customer in southern Denmark" is not.
Deferring the decision. While management deliberates, employees use a free account. No decision is also a decision, just an uncontrolled one.
Starting with the hard case. The fastest route to a decision is to try it on something harmless.
The first morning
- Inventory, one hour. Write down your ten most important data sources.
- Classification, one hour. Give each source a level. Disagree out loud, that is where the work is.
- Rule, ten minutes. Write one sentence: RESTRICTED never leaves the building.
- Vendor email, ten minutes. Send the three questions.
- Try it, thirty minutes. Run an analysis on pure PUBLIC data. Competitors' published accounts, market trends. No personal data at all.
Point 5 is what moves things. Once management has seen what the system does with data nobody is nervous about, the question of INTERNAL becomes concrete rather than theoretical.
What the board should ask
If the subject reaches a board meeting, three questions establish whether the vendor has thought it through: where does our data physically sit, can we see what the system built its answer on, and may we choose something other than the recommendation.
Those three demands, and how to hear from the answer whether any thought went into it, are in three things your board should demand.
So the answer to the original question is yes. You can use AI on your own data. You just need to have decided which data before somebody asks.