Copyright and Generative AI: Proving the Origin of Datasets with V-PROOF
Copyright and Generative AI: Proof of Origin as a Legal Defense
EU Directive 2019/790 and Article 53 of the EU AI Act converge on the same point: without cryptographic traceability of the training data, there is no possible legal defense.
Key Findings
- More than 30 active lawsuits against LLM models in the U.S. and Europe (NYT vs. OpenAI, Getty vs. Stability AI, Universal Music vs. Anthropic). The common thread: What data was used, and with what authorization?
- EU Directive 2019/790, Article 4, allows for an exception regarding text mining but recognizes the rights holder's right to "expressly reserve" the right to use the data, thereby creating a retroactive legal liability for models that have already been trained.
- EU AI Act, Art. 53(1)(d), requires GPAI providers to publish a "sufficiently detailed summary" of the training data, including identification of sources and licenses.
- Without cryptographic " hash " of the dataset from the time of ingestion, it is technically impossible to prove exactly which version of the data was used, when it was used, and whether the data subject's opt-out request had already been recorded.
- V-PROOF Generates an SHA-256 hash of the entire dataset at the time of ingestion, which is timestamped in Base L2 before the first training cycle, creating evidence prior to the process, not after it.
Two standards, the same problem with evidence
The training of generative AI models is under legal scrutiny on both sides of the Atlantic. In Europe, the framework is built on two mutually reinforcing pillars.
Why Internal Logs Are Not Enough
The typical response of organizations to an audit is to submit system logs, contracts with data providers, and internal compliance statements. None of these elements can withstand rigorous adversarial analysis for three fundamental reasons.
Internal logs can be modified by the system operator. They do not have an external timestamp. They do not generate a verifi hash. In the event of a dispute, the other party may question their integrity without any technical evidence to refute it.
The second loophole is more subtle: the time of capture. An opt-out notice published by a rights holder on March 15 is legally valid if the model began training on March 20. But if the system cannot accurately demonstrate when it processed the specific data, the opt-out applies by default.
The third gap concerns the exact version of the dataset. Datasets evolve: they are cleaned, filtered, and new sources are added. Without version- hash, it is not possible to prove that the model was trained on the cleaned, post-filtering version, rather than on the previous version that contained protected content.
"Can you verify which exact version of the dataset was used, the exact time it was processed, and that at that time there were no opt-out requests on record from the data subjects included?" · If the answer cannot be verified by an independent third party, the legal risk remains unresolved.
Origin Evidence Pipeline, V-PROOF
V-PROOF It generates cryptographic evidence of the dataset at the time of ingestion, prior to any training process. The chain is immutable and verifiable by third parties without the need for access to internal systems.
12.8 GB · 4.2M documents
with no middleman
The result is a dataset “V-SEAL ”: a record that includes the exact version of the dataset, the complete SHA-256 “ hash,” the IPFS CID, the block number in Base L2, and the precise “ timestamp ” of the operation. This record can be verified by any auditor or court without requiring access to the organization’s internal systems.
Covered Articles, EU AI Act & Copyright Policy
| Article | Obligation | Evidence V-PROOF | Modules |
|---|---|---|---|
| EU Directive 2019/790 · Copyright in the Digital Single Market | |||
| Article 3 | Text mining for research: demonstrating that the use falls within the exception | ✓ Hash of the dataset + source classification upon ingestion | SHA-256 IPFS |
| Article 4 | Commercial use: demonstrate that there was no opt-out option at the time of processing | ✓ The timestamp records the exact time, either before or after the exclusion | L2 Base V-SEAL |
| EU AI Act · GPAI models (effective August 2025) | |||
| Art. 53(1)(b) | Technical documentation for the training datasets used | ✓ Complete record: version, sources, hash, timestamp | Governance IPFS |
| Art. 53(1)(c) | Policies Implemented to Comply with Copyright Laws During Training | ~ V-PROOF documents the process; the exclusion policy is set by the operator | Governance |
| Art. 53(1)(d) | Publish a sufficiently detailed summary of training data (sources + licenses) | ✓ Verifiable export of the dataset record with license metadata | V-SEAL API |
| Art. 55(1)(a) | GPAI Models of Systemic Risk: Risk Assessment Including a Dataset | ✓ Full traceability for audits by the EU Intellectual Property Office | L2 Base V-SEAL |
Why V-PROOF is the correct technical answer
Key advantages of the cryptographic approach outlined in * V-PROOF * compared to declarative compliance approaches.
Does your GPAI model include evidence of the dataset's origin?
V-PROOF 's Strategic Assessment evaluates your legal Exposure at EU AI Act under Article 53 and the Copyright Directive, and outlines a plan for implementing cryptographic evidence within 4 weeks.
Request a Diagnosis →