EU AI Act training data summary: what it means if you license your company's data
The EU AI Act training data summary is a public document that general-purpose AI model providers must publish, using a Commission template, and licensed datasets are one section of it. The duty is the provider's, not the supplier's, but buyers may ask you for provenance, rights and redaction details to describe your dataset.
What does the EU AI Act training data summary mean for a company that licenses data?
The EU AI Act requires providers of general-purpose AI models to publish a summary of the content used to train them, and the licensed datasets they buy are one section of that summary. For a company that holds business records, the practical result is that an AI developer may need to describe your dataset in general terms, and may ask you for provenance and rights information to do so.
The duty sits on the model provider, not on you. Article 53(1)(d) of the AI Act obliges providers to draw up and make publicly available a sufficiently detailed summary of the content used for training, following a template from the EU AI Office. The Commission has since published that template, so providers now work from a standard form. Application dates in the Act have been amended since adoption, so check the consolidated text for the dates that apply today.
This page is written for owners and executives of US companies who are considering a data license. It is not for model providers.
What does the template ask a provider to disclose?
The template groups disclosures by where the training content came from, then asks how it was processed. The headings below are a simplified reader's guide and may not match the template's current wording or fields, so read the Commission's published template itself before relying on any row.
| Template area | What a provider describes | What it can mean for you as a supplier |
|---|---|---|
| General information | Who the provider is, the model, and the types of data used | Nothing is asked of you directly |
| Publicly available datasets | Large datasets that anyone can download | Not relevant to a private license |
| Private datasets | Licensed or otherwise non-public datasets obtained from third parties | This is where your dataset would sit |
| Crawled or scraped content | Content collected from the web | Not relevant unless your public content was scraped |
| User data and synthetic data | Data from the provider's own users, and data generated by models | Not relevant to a private license |
| Data processing | Steps such as respecting rights reservations and removing unwanted content | A buyer may ask how you handled sensitive material |
The key point for suppliers is the private datasets area. A provider is asked to describe such datasets in general terms. The template is not designed to make a provider publish your records, and it is not a request for the contents of your files.
Will my company be named or my records published?
Not by the template itself, but expect the buyer to ask you for enough information to describe the dataset accurately. A summary is meant to give rightsholders and the public a general picture of training content, not to reproduce it.
What you should expect to be asked, whichever buyer you work with:
- What kinds of records are in the dataset, in plain terms (for example support tickets, internal documents or project files).
- What time period and which systems they came from.
- Whether you created the records and have the right to license them.
- Whether personal data, client data or other sensitive material was removed or redacted first.
- Whether any of it is already licensed to someone else.
Whether a supplier is named in a public summary is a commercial and contractual question. Settle it in the license: ask for a clause on how the dataset may be described, what may be disclosed, and whether your company name can appear. Nothing is binding on you until you agree price and terms and sign.
How does a US company fit into an EU rule?
The AI Act reaches providers who place models on the EU market, wherever they are based, so a US supplier can be pulled into the paperwork indirectly through a buyer's obligations. Personal data of people in the EU adds a second layer: the GDPR can apply to US organizations that offer goods or services to, or monitor the behavior of, people in the EU.
| Situation | What to check | Typical outcome to confirm with counsel |
|---|---|---|
| Records contain only US business activity | Whether any customer or staff data relates to people in the EU | Often limited EU exposure, but verify rather than assume |
| Records include EU customers or staff | Whether GDPR applies to your processing and what redaction is needed | Redaction or exclusion agreed before delivery |
| Buyer is an EU-facing provider | What provenance statements the buyer wants in the contract | Representations about rights and origin, within what you can honestly state |
| You are asked to warrant that no copyrighted third-party material is present | Whether your records include client or vendor material you do not own | Narrower warranty, or excluded folders |
| Buyer wants to describe the dataset publicly | Naming, wording and approval rights | A disclosure clause you have reviewed |
This is general information, not legal, tax or financial advice. Confirm with your own counsel before acting, especially on questions about EU personal data.
What should I prepare so a buyer can describe my dataset?
Prepare the facts a provider would need, before anyone asks. This is the same groundwork a licensing review needs anyway.
- A one-page description of each system in scope, the years it covers and who created the records.
- A record of which folders, channels or ticket queues are excluded and why; the default exclusion list is a starting point.
- Confirmation that you own the records, or have consent where clients or partners are involved.
- A view on personal data: what categories exist, which are removed, and whether health information is present (see HIPAA and AI training data).
- Any employee notices or policies that cover internal communications, such as the employee privacy notice template for California.
- A statement of any prior license of the same records.
- A named person with authority to approve the description and disclosure terms.
Consent and documented rights are the foundation of the whole exercise, which is why consent matters in AI data licensing.
How does this fit with US state rules?
The EU summary is one of several transparency regimes, and US states are adding their own. Some state laws touch training data disclosure and consumer notice in ways that differ from the EU approach. The overview of state AI laws in 2026 covers the ones to watch. Keep the two apart: the AI Act template is about what a model provider publishes, while state laws may place duties on the companies that hold or use the data.
How does SourceX handle this?
SourceX manages data licensing for companies, from sourcing and rights review to delivery and payment. You keep ownership; data is licensed, not sold. De-identification and redaction requirements are agreed with the company before any work begins, and data is delivered only after an executed agreement and the company's authorization. Partners never handle records: a partner only makes the introduction.
Wording on how a dataset may be described to the public belongs in the agreed terms, so raise it early with the SourceX team when you complete the data inventory. See how it works for the full process.
Who does this not apply to, and when should you wait?
Pause before licensing, or fix the gap first, if any of these is true:
- The records mainly belong to your clients and they have not agreed.
- The data is mostly consumer personal information, or health information without HIPAA authorization or de-identification.
- You cannot say who created the records or what they contain.
- The same data is already licensed for AI training.
- A court, trustee or assignee controls the assets and has not been involved.
A company that cannot answer the provenance questions today can often build the answers over a few weeks. The company fit checker gives a preliminary, non-binding screen with no contact details required.
Next step
If you advise or know a US company with 50+ full-time employees at peak (contractors excluded) that keeps years of records across many systems, register as a partner and introduce it. Partners earn 25% of the eligible platform fees SourceX actually collects, capped at $100,000 per referred company, paid only after the buyer pays and SourceX receives its fee; no reward is guaranteed. Company owners can also apply directly at sourcex.si/apply.
- Step 1Share your linkSend your personal link to a company you know.
- Step 2Company appliesThe company applies itself at /apply.
- Step 3Buyer selects and paysThe buyer selects and pays for the data and SourceX receives its fee.
- Step 4You get your rewardYour share of SourceX fees becomes payable.
Common questions
Do I have to publish anything under the EU AI Act if I license my data?
No. The publication duty falls on the provider of the general-purpose AI model, not on the company supplying a licensed dataset. You may be asked for provenance and rights information so the provider can describe your dataset accurately, and the license should say what may be disclosed.
Can a buyer put my company name in its public training data summary?
That depends on the contract. The template asks providers to describe private datasets in general terms and does not itself require naming suppliers. Treat naming and wording as a negotiated point: ask for a clause covering how the dataset is described and whether your name can appear.
Does the training data summary reveal my actual records?
No. The summary is meant to give rightsholders and the public a general picture of training content. It is not a copy of the data, and it is not a request for your files. Your records are delivered only to the buyer, under the signed agreement.
Does the AI Act apply to a US company that only licenses data?
Not directly as a supplier. The Act reaches model providers placing models on the EU market. A US supplier is affected indirectly through buyer requests. If your records contain personal data of people in the EU, the GDPR may be a separate question to take to counsel.
What should I check first before a buyer asks about provenance?
Confirm that your company created the records, list the systems and years covered, note what is excluded, and check for client, health or employee personal data. Then name one executive who can approve the dataset description. These steps also feed the licensing data inventory.
Related pages
- What data should be excluded from AI training? A default exclusion list
- HIPAA and AI training data: what the rules allow and what stays out of a license
- Employee privacy notice template for California employers, clause by clause
- Why consent is the foundation of AI data licensing
- State AI laws in 2026 that touch AI training data: what to check before licensing
- How SourceX US company data referrals work
Free resources
- Due diligence checklist generator — A tailored document request list by deal type.
- Cash flow calculator — A 12-month cash forecast with shortfalls highlighted.
- Referral earnings calculator — Hypothetical partner earnings with the per-company cap.
- All free tools · MCP resource center
By SourceX Partnerships Team · Published 2026-10-09 · Updated 2026-10-09
Know a US company with valuable proprietary data?
Become a referral partner from anywhere we support, get your link and introduce an owner or authorized decision-maker.
Refer a company →I own a business
Explore licensing your company's data to AI developers worldwide. Start a short assessment; no uploads needed.
Start an assessment