AI: can you really keep control of the data you share?
Quick reply
Shadow AI, the GDPR and the AI Act: what AI tools really do with your files, and the levers for keeping control of your company data.

Picture this: an HR manager summarises payslips in ChatGPT to save time. A salesperson pastes a contract into Gemini to pull out the key clauses. An engineer submits their source code to an AI tool to fix a bug.
None of these people consulted the IT department. None of them wondered where that data was going. That is precisely where the problem starts. The real question is not "can we use AI in our company?", it is "do we actually know what we are giving it to read, and what it does with it?"
The problem often begins well before that: when a company's sensitive documents are scattered across mailboxes, servers, public clouds and collaborative spaces, with no real document governance.
What AI actually does with your data
Not every AI tool behaves the same way with your data. The fundamental distinction lies between two very different architectures.
"Classic" LLMs (such as the free version of ChatGPT) are trained on vast corpora of data and can, by default, use your conversations to improve their models. RAG LLMs (Retrieval-Augmented Generation), by contrast, reach your files and document stores in real time, which raises specific questions about access control.
What OpenAI says about its various plans is instructive. For the free version of ChatGPT, conversations are kept indefinitely and can be used to train the model, unless the user turns that option off manually. The data is hosted in the United States, with transfers outside the EU covered by standard contractual clauses, which some European authorities consider insufficient.
For OpenAI's Business and Enterprise plans, the situation is different: your data is not used to train the models, a DPA (data processing agreement) is available, and SOC 2 type 2 and ISO 27001 certifications have been obtained.
Shadow AI: the risk IT departments underestimate
Shadow AI means employees using artificial intelligence tools without authorisation or oversight from management or the IT department. Opening ChatGPT in a browser takes three seconds. Waiting for the IT department to deploy an official AI tool can take months.
The figures speak for themselves:
- 68% of employees use unauthorised AI tools at work, against 41% in 2023 (Gartner, study of 500 companies, 2025).
- 61% of executives themselves use generative AI through personal accounts at least once a week (Microsoft France / YouGov, January 2026).
- 33% of employees admit to having shared sensitive data (employee records, financial information, research) in unapproved tools (BlackFog, 2026).
- The average additional cost per incident tied to Shadow AI reaches 670,000 dollars compared with other security incidents (IBM, cost of a data breach report, 2025).
What makes Shadow AI more dangerous than classic shadow IT is how invisible it is. Pasting a document into a chatbot triggers no alert in the usual monitoring systems. Without a clear framework, these small actions multiply and escape any control.
What the regulations say in 2026
Using AI in business is governed by two complementary European regulations: the GDPR, which covers personal data, and the AI Act, which covers AI systems themselves. Neither replaces the other. They apply at the same time.
The GDPR applies as soon as personal data is involved
As soon as you send personal data (a client's name, an email address, an HR record) to a third-party AI tool, you are the data controller. That brings a number of obligations:
- Signing a DPA (data processing agreement) with the AI provider.
- Recording the processing in your register of activities (article 30 of the GDPR).
- Carrying out a data protection impact assessment (DPIA) where the processing presents a high risk, which the CNIL recommends as soon as sensitive data is processed at scale.
- Checking that hosting takes place on European soil, or with sufficient transfer safeguards.
The AI Act: a phased timetable with severe penalties
The AI Act has come into force in stages. Since February 2025, the outright prohibitions have applied (social scoring systems, behavioural manipulation). Since August 2025, the obligations for general purpose AI models (GPAI) such as GPT or Claude have been in effect. From 2 August 2026, the majority of obligations apply, particularly for high-risk systems.
The regulation classifies AI systems across four levels of risk. For companies, the "high risk" category covers everyday uses directly: sorting CVs, customer scoring, employee assessment and decision-making systems in financial services.
The AI Act also imposes an AI literacy obligation (article 4), in force since February 2025: everyone who deploys or supervises AI systems has to have a sufficient level of competence. That applies to your IT department, your managers and your teams working with these tools.
Can you really keep control of the data you share with AI?
Yes, under certain conditions
Several levers let you keep a hand on your data:
- Choose tools whose contractual terms guarantee that data will not be reused to train models, with a signed DPA.
- Opt for solutions hosted in Europe, to free yourself from the risks tied to the American Cloud Act, which allows the American authorities to demand access to data held by companies under American law, whatever country it is stored in.
- Deploy models in a private or on-premise environment for the most sensitive uses, so that data does not pass through third-party servers.
- Put fine-grained access segmentation in place before connecting any AI tool to your document systems.
No, if you do nothing
Without an internal AI policy, data moves around freely. An AI agent connected to your CRM, your mailboxes and your shared folders reaches all of those resources, and can become a prime entry point for an attacker if it is compromised. Access segmentation is not optional: it is the precondition for any responsible AI deployment.
Controlling AI data starts with controlling data, full stop
Before asking what a model does with your data, you have to be sure you yourself have a clear, controlled view of your own documents.
That means knowing precisely where your sensitive files are stored, who has access to them and since when, what has been done to those documents, and in which countries the data is hosted.
That is exactly what sovereign document management solutions allow: centralise, control and trace before you connect. NetExplorer, hosted entirely in France, lets you set granular access rights document by document, follow every access and change, and share files securely with third parties without that data leaving French soil.
Before connecting AI to your documents, the first step is therefore to know exactly what you are giving it access to. That is a question of governance, not of technology.
Conclusion
Controlling the data you share with AI is possible, but it takes a deliberate approach, not simply a purchasing decision. Document security in the age of AI rests on three fundamentals: knowing where your data is, controlling who reaches it, and choosing technology partners who let you stay in charge of it.
The regulatory framework, the GDPR and the AI Act, now provides a floor of requirements. But companies that settle for the regulatory minimum are taking risks the law does not yet cover in full. Shadow AI, for its part, is not waiting.
Other articles you might like
Strengthening your resilience against cyberattacks
Ransomware and human error: the technical and organisational levers that build resilience and let you react quickly after an incident.


But let's be honest, our cloud-based file storage and sharing solution is much easier.


