Real documents versus a safe exercise is a genuine trade off
People learning to use an AI tool pick it up fastest on material that resembles what they actually do every day. A generic exercise text written just for the lesson tends to feel disconnected, and what people learn transfers less well back into their actual work. That does not make the choice simple, though: a real company document nearly always contains something the company would not want shared freely, whether that is employee or customer personal data, commercial terms, or internal figures. The moment that document appears in a training environment, particularly if someone uploads it into a third party AI tool, the risk stops being theoretical.
Deciding whether and how to bring real documents into training is not a question of convenience in preparation. It weighs two genuine goods against each other: how well the skill transfers into real work, and how well the data entering the training is protected. The right answer also differs document by document, not just company by company.
Why real material genuinely helps
The format and structure of real documents, the length people actually deal with, the ambiguities and phrasing that show up in practice, are exactly what a generic sample text tends to lack. Someone who practises on a document resembling what lands in their inbox on a Monday morning remembers more easily how to phrase a request and what to check in the output. That advantage is the main reason companies lean towards real documents in the first place, which is why this article does not recommend excluding real data from training as a blanket rule.
What real material genuinely risks
The risk has two distinct layers, and it helps to keep them apart. The first is exposure to people inside the training session who would not otherwise have access to the original document, a colleague from another department, for instance. The second, usually the more serious layer, is exposure to the AI vendor’s own service if someone uploads the document into the tool during the exercise. GDPR Articles 5 and 6 require a clearly defined purpose and legal basis for processing personal data; a training exercise does not automatically create that basis just because it is useful to the company. Article 32 then requires appropriate security for the processing, which covers who the data reaches and where it travels during training. EDPB Opinion 28/2024 addresses when personal data contained in material used for AI models may be further processed, and confirms that the assessment always turns on the specific circumstances of the processing, not on a general assurance that it is ‘only training’.
Four middle paths between fully real and purely artificial
Between the extreme of ‘we will use actual customer contracts’ and the extreme of ‘we will only use invented text’ sit four practical options.
A redacted excerpt takes a real document and removes or blacks out the sensitive parts, names, numbers, identifiers, leaving only what the exercise needs. This route is fast but needs care: a poorly done redaction can leave sensitive detail in metadata or allow it to be inferred from what remains.
A synthetic document is built from scratch to closely match the structure, length and typical phrasing of a real document of that type, without containing any actual data. It takes time to prepare but effectively removes the leak risk, because there is nothing genuine in it to leak.
A real document used only in an approved company environment means the document stays genuine but is processed exclusively inside a tool the company has approved for that purpose, typically its own tenant with contractual data handling terms, not a public version of the tool.
A real document used only for reading, never uploaded into any tool, means the document serves as a visual or printed example that participants view and discuss, but nobody uploads or pastes it into an AI interface at any point.
Decision table by document class
The table below is limited to what is worth considering for a given class of document inside the training session itself. It does not attempt a general classification of company data outside that context. The ‘decision’ for every class always belongs to the data protection officer or lawyer, not the trainer.
| Document class | Redacted excerpt | Synthetic document | Real document in an approved company environment | Real document, read only, never uploaded |
|---|---|---|---|---|
| Publicly available company material (press releases, public website) | worth considering, usually with no changes needed | worth considering if it serves the purpose equally well | worth considering | worth considering |
| Internal document with no personal data and no trade secrets | worth considering | worth considering | worth considering after the data owner approves | worth considering |
| Document with employee or customer personal data | worth considering only after the data protection officer confirms redaction is sufficient | preferred option if the structure is enough for the exercise | worth considering only after a separate legal review and contractual terms | worth considering after data protection officer review |
| Document with trade secrets or sensitive contract terms | worth considering only after a lawyer confirms the excerpt does not disclose the secret | preferred option | worth considering only after a separate legal review and contractual terms | worth considering after legal review |
| Document with special confidentiality (for example privileged legal correspondence, health data) | not recommended without explicit approval from both lawyer and data protection officer | preferred option | worth considering only in an exceptional, separately approved case | not recommended without explicit approval |
‘Worth considering’ in the table is not automatic permission; it means the option is on the table in principle, and only the data protection officer or lawyer decides whether a specific document actually meets the conditions. Combining several documents together, or stripping obvious identifiers, does not by itself amount to anonymous data in the legal sense of the term; only the data protection officer or lawyer can reach that conclusion, based on a specific review, not the person preparing the training material.
Where this decision sits in the rest of training planning
The choice among these four paths should be made before the training outline is built, because it shapes how much preparation time is needed and who has to sign off on the material beforehand. How this decision connects to the broader shape of a training programme, including whether to run something general or built around your own industry, is covered in general or industry specific AI training, and common ways preparation goes wrong are covered in the mistakes that ruin company AI training. Whether employees may routinely put company data into publicly available AI tools outside a training context is a separate question, covered in is it safe to put company data into ChatGPT. Who actually sits in the room handling this material also matters; see who to pick for the first AI training pilot group.
Sources and limits
Articles 5, 6 and 32 of GDPR support the claims about purpose, legal basis and security of processing, but none of them generates an automatic answer for a specific company document; that is always an individual review. EDPB Opinion 28/2024 concerns the processing of personal data in connection with AI models generally, and is cited here only for the principle that any assessment must be tied to specific circumstances, not as a how to guide for a specific training procedure. NIST AI 600-1 describes a risk profile for generative AI, including risks tied to data entering such systems; it is a general framework, not a legal opinion for the Czech or European context. This article does not claim that a redacted excerpt or a synthetic document is always sufficient, nor that aggregated or blacked out data is automatically anonymous; both require review by the data protection officer and company lawyer for each specific document.
Frequently asked questions
Is it safe to use a training document where we only removed names?
Removing names alone does not necessarily make a document safe or anonymous; remaining context, numbers or a combination of details can still allow a person or a company to be identified. Whether a specific edit is sufficient is a judgement for the data protection officer or company lawyer, and this article does not substitute for that review.
Can we just display a real document on screen without uploading it into an AI tool?
Yes, that is one of the four middle paths described above. The document serves as a visual or printed example of structure and format, participants view and discuss it, but nobody uploads or pastes it into any AI interface, so the sensitive content never leaves the company environment toward the tool's vendor.
Will a synthetic document work, or will participants notice it is artificial and stop taking the exercise seriously?
It depends on how carefully it is prepared. A synthetic document that faithfully copies the real structure, typical length, phrasing and common errors of that document type usually works about as well as a genuine one. A rough, unrealistic mock up noticeably reduces the value of the exercise, which is why the time spent preparing it properly tends to pay off.
Who inside the company decides which class a document falls into?
This article cannot make that call on its own, and neither should a trainer. Classifying a specific document and deciding whether it can be used in a given training format is the job of the data protection officer, working with the company lawyer and the data owner, not the person running the session.
Do the same rules apply when an external provider runs the training?
The risk is usually higher with an external provider, because the document leaves the company environment for a third party the company may not trust to the same degree as its own staff. Using real documents with an external provider therefore calls for closer review of the vendor, the data processing terms, and often prior approval from the data protection officer.
SOURCES AND VERIFICATION
reviewed Lukáš Dlouhý ·