AI Risk & Contractual Liability

Your AI vendor said the training data was properly licensed

The FTC concluded three enforcement actions against AI model vendors who misrepresented training data sourcing as licensed, public domain, or fair-use qualified, when the data actually came from unauthorised web scraping and copyrighted material used without consent.

Eliga Consultancy 19 August 2026 AI Risk, Contract Law 8 minute read

The U.S. Federal Trade Commission concluded enforcement actions against three AI model vendors on 19 August 2026. Vendor representations about training data are now enforceable and auditable, and customers who continue using a model after notice of infringing content carry their own liability.

01

Three enforcement actions

Core enforcement theory

Vendors’ misrepresentation of training data sourcing is an unfair and deceptive practice. Customers rely on these representations, and proving them false is grounds for FTC enforcement.

The Federal Trade Commission announced three enforcement actions against major AI model vendors on 19 August 2026, concluding settlements involving a major generative AI provider, an image generation platform, and an industry-specific model provider.

In each case, the FTC investigation found that the vendor had publicly stated, in marketing materials, documentation, or customer disclosures, that training data was properly licensed or in the public domain. Investigation revealed substantial portions of training data actually came from unauthorised web scraping and included copyrighted material used without copyright holder permission.

The FTC’s theory: this is an unfair and deceptive trade practice under the FTC Act. Vendors made material misrepresentations about product characteristics that would affect customer purchasing decisions and liability exposure. A customer who chose a vendor because it claimed ethically sourced training data, when that data was actually infringing, was deceived.

02

What vendors falsely claimed

Three categories of claim

Licensing, public domain, and fair use arguments. FTC investigation found each was materially false.

“Data is properly licensed.” Vendors claimed training data was licensed from content providers with permission. In reality, 30 to 40 per cent of training data in the investigated cases came from unauthorised web scraping, and licenses had not been obtained for most of it.

“Data is public domain or open-source.” Vendors claimed training data was sourced from public domain and open-source repositories where use is permitted. In reality, vendors had scraped copyrighted news articles, academic papers and creative works from public websites. Presence on a public website does not make content public domain.

“Use qualifies as fair use.” Vendors claimed transformative AI training constitutes fair use under U.S. copyright law. In reality, vendors extracted verbatim sections of copyrighted text and code in model outputs. That is reproduction, not transformation, and the fair use defence failed.

03

What actually happened

Vendors misrepresented not just the legality of data sourcing, but the scope of infringing content. Claiming “properly licensed” created an expectation of minimised infringement that investigation showed was the opposite of reality.

In each case, FTC investigation revealed a consistent pattern: automated web scraping without permission from content owners, no verification of whether scraped content was licensed or public domain, substantial copyrighted material folded into training data including news articles, academic papers, published books, photographs, creative writing and source code, no notification to copyright holders before use, and in some cases model outputs reproducing training data nearly verbatim.

Copyright holders discovered the use of their work only when it appeared in model outputs, not through any notification from the vendor.

04

Settlement requirements

Enforcement teeth

Settlements are not just fines. Vendors face ongoing third-party audits, detailed record-keeping, and mandatory customer notification.

The three settlements impose ongoing obligations on vendors:

  • Third-party audit of training data provenance: independent auditors must certify or refuse to certify that data was obtained with appropriate permissions, annually for five years.
  • Data inventory and provenance records: detailed records showing the source of each training data component, licensing status, and whether it includes copyrighted material.
  • Identification and notification of infringing content: vendors must analyse training data to identify copyrighted material used without permission and notify the copyright holders.
  • Customer notification: vendors must notify all customers who purchased models based on the misrepresented training data, disclosing that infringing content was included and that FTC enforcement concluded.
  • Refunds or termination rights: customers who relied on training data representations may be entitled to a refund or to terminate the contract without penalty.
05

Customer liability after notice

A game-changer for customer liability

Before settlement, a customer could argue reliance on vendor representations. After notification of infringing content, continuing use exposes the customer to its own copyright liability, not just the vendor’s.

The settlements require vendors to actively identify which training data is copyrighted and notify affected customers. That means vendors will soon be sending notices along the lines of: your model was trained on data that included copyrighted material from named publishers or authors, and recommending customers audit model outputs for verbatim copying, review indemnity provisions, and take legal advice on liability exposure.

Historically, customers could argue they relied on vendor representations and the vendor bore the risk. After these settlements, customers have explicit notice that infringing content is in the model, and continuing to use it after that notice is a decision the customer now owns.

06

What to demand from vendors

Template language

“Vendor warrants training data sourcing is lawful and fully licensed. Customer has audit rights. Vendor indemnifies Customer against FTC enforcement and copyright claims arising from vendor’s training data practices.”

If you license AI models from vendors, your agreements need three sections:

  • Training data warranty: an unqualified warranty that training data was obtained with appropriate permissions and does not include substantial copyrighted material obtained without permission, with refund or termination rights if breached. Push back on vendor attempts to qualify this with “to the best of vendor’s knowledge.”
  • Audit rights: the right to conduct or commission independent audits of training data provenance and copyright compliance, with vendor cooperation and documentation. Push for reasonable audits at any time, or annual audits at the vendor’s expense.
  • Indemnity for FTC liability and copyright claims: indemnity against any FTC enforcement action or settlement arising from the vendor’s training data practices, and against copyright infringement claims arising from the customer’s use of the model where the infringing content sits in the vendor’s training data.

Review your existing vendor agreements now. If one says “we cannot warrant training data sourcing,” treat that clause as material risk and ask the vendor to either provide training data warranties or agree to refund or termination if FTC enforcement later concludes the training data is infringing.

Questions this raises

Six questions the FTC enforcement actions tend to prompt, answered directly.

Which AI vendors did the FTC take action against?

Three major AI vendors concluded settlements with the FTC for misrepresenting training data sourcing. Case details are public in the FTC enforcement announcements, though this piece does not name the specific vendors.

What does the settlement require of AI vendors?

Five years of independent third-party audits of training data provenance, maintained data inventory records, notification to copyright holders whose work was used, and notification to affected customers, with compliance beginning within 90 days of settlement.

Can vendors continue selling the affected models?

Yes. The settlements do not require model withdrawal. Vendors must disclose infringing content in the training data and notify customers, but can continue offering the model.

Are customers liable if their AI system used infringing training data?

Liability depends on notice. Before disclosure, a customer could reasonably rely on vendor representations. After notification that infringing content is present, continuing to use the model exposes the customer to its own copyright liability.

What should a business audit before licensing an AI model?

Request the vendor’s training data documentation, commission an independent audit of licensing claims, and check specifically for known copyrighted works such as books, news articles and academic papers inside the training data.

What should be demanded in AI vendor contracts?

An unqualified warranty that training data is properly licensed, audit rights, indemnity against FTC enforcement and copyright claims, and refund or termination rights if the warranty is breached.

Sources

  1. U.S. Federal Trade Commission, Enforcement Actions on AI Training Data Misrepresentation, 19 August 2026.
  2. FTC, Settlement Agreements with AI Model Vendors, August 2026.
  3. FTC, Third-Party Audit Requirements for Training Data Provenance, August 2026.

Auditing your AI vendor agreements for training data risk?

Eliga provides embedded commercial and technology counsel to scaling businesses. FTC enforcement shows training data sourcing claims are now enforceable and auditable. We help organisations review AI vendor agreements, audit training data sourcing, and renegotiate contracts to include training data warranties and audit rights.

This page is general information about US and UK commercial and technology law. It is not legal advice and does not create a solicitor-client relationship. Take specific advice on anything you are about to negotiate or sign.