Rights & provenance

The EU AI Act and training data

Last updated 2026-07-22

The EU AI Act — Regulation (EU) 2024/1689 — is the first comprehensive AI statute, and the first to demand public disclosure about training data. Its general-purpose AI chapter has applied since August 2, 2025, and Article 53 puts two obligations on every provider of a general-purpose model placed on the EU market. First, the provider must maintain a policy to comply with EU copyright law, including identifying and honoring rights reservations expressed under Article 4(3) of the DSM Directive — the machine-readable opt-outs owners use against text-and-data mining. Second, it must publish a "sufficiently detailed" summary of the content used to train the model, in the mandatory template the Commission’s AI Office adopted on July 24, 2025. The obligations sit on the model provider, not the data supplier. But they reshape what both sides of a data deal need on paper, because a provider can only disclose and defend what its suppliers documented.

Why it matters to a buyer

The common misreading is that this is a European problem a US-trained model can ignore. It is not. The Act applies to general-purpose models placed on the EU market regardless of where the training ran; if your model is available in the EU, Article 53 reaches it. The timing is the point: providers have been under the obligations since August 2, 2025, and from August 2, 2026 the Commission can enforce them directly, with fines up to 3% of global annual turnover or €15 million, whichever is higher. That converts training-data documentation from a diligence preference into regulatory exposure. Note also what the summary is: public. Rights holders, competitors, and plaintiffs’ counsel will read it. The template asks for your main sources by category — licensed data, public datasets, scraped content, user data, synthetic data — with indicative size ranges, plus a description of how copyright reservations were respected. A corpus you cannot describe is a corpus you cannot lawfully ship in Europe, and a copyright policy that cannot show how opt-outs were honored is an open finding waiting for a complaint. This is why documentation demands now flow down the supply chain: a buyer needs each supplier to state in writing what the material is, where it came from, and that training rights were granted — because the buyer must summarize exactly that in public and stand behind it to the AI Office on request.

Why it matters to a data owner

For owners, the Act does two distinct things. It gives the EU opt-out teeth: every GPAI provider’s copyright policy must honor a properly expressed Article 4(3) reservation, so reserving your rights is no longer a gesture crawlers can ignore at no cost. But a reservation only stops use — it does not pay you. The second effect is the commercial one. Because buyers must now account publicly for what they trained on, data that arrives with a licence, provenance, and consent records is worth more than data that arrives with a shrug. A supplier who can hand a buyer the artifact its Article 53 summary needs — a written grant of training rights and a documented source — is selling compliance, not just content. Expect EU-facing buyers to add cooperation clauses to data licences: obligations to describe the material accurately, to disclose how it was collected, and to stand behind that description. Suppliers who prepare the documentation before the ask close faster and on better terms.

Current legal status

Regulation (EU) 2024/1689 applies in stages. The GPAI obligations in Article 53 have applied since August 2, 2025: Article 53(1)(c) requires a copyright-compliance policy that honors rights reservations under Article 4(3) of Directive (EU) 2019/790, and Article 53(1)(d) requires a publicly available summary of training content using the template the Commission adopted on July 24, 2025 — covering provider and model identification, main data sources by category with indicative size ranges, and how copyright, illegal content, and data protection were handled. The voluntary General-Purpose AI Code of Practice, published July 10, 2025 and confirmed adequate by the Commission and the AI Board on August 1, 2025, is the recognized compliance path; its copyright chapter commits signatories to lawful data sourcing, respect for machine-readable opt-outs, safeguards against infringing outputs, and a complaints channel for rights holders. OpenAI, Google, Anthropic, Microsoft, Amazon, and Mistral signed; xAI signed the safety chapter only; Meta declined, citing legal uncertainty — and non-signatories still owe the underlying statutory obligations, just without the Code’s presumption of good faith. Enforcement is phased: the Commission’s supervisory and fining powers over GPAI providers begin August 2, 2026 — penalties up to 3% of global annual turnover or €15 million — and models placed on the market before August 2, 2025 have until August 2, 2027 to comply. One 2026 development matters here mostly for what it did not change: the Digital Omnibus package, agreed politically in May 2026 and endorsed by the European Parliament on June 16, 2026, postpones the separate high-risk system deadlines into 2027 and 2028, but left the GPAI training-data obligations and the August 2026 enforcement start in place.

What fiund does about it

fiund’s model produces the paperwork Article 53 assumes exists: a signed licence granting training rights, provenance records identifying the source, and consent artifacts where people are identifiable — the inputs a buyer’s public training-content summary and copyright policy are built from. Licensed acquisition is the simplest copyright policy there is.

Sources

← All rights & provenance guides

More on rights & provenance

Want data that clears this in diligence?

Whether you're building a model or sitting on an archive, the first conversation is short and specific.

Send a brief