The publishing industry's AI copyright settlements are setting the price of a training dataset
The wave of settlements between major AI labs and groups representing authors, publishers and news organisations over unauthorised use of copyrighted material in training datasets has begun to establish something the industry lacked entirely before: an approximate market price for licensed training data, expressed through the per-work or lump-sum compensation figures that have emerged from settled cases and voluntary licensing agreements alike. These figures, while modest relative to the scale of AI-lab revenues, have created a template that licensing negotiators on both sides now reference when structuring new deals, moving the industry gradually from an unlicensed-scraping default toward a negotiated-licensing norm, at least for content published after the major legal disputes began. The unresolved and more consequential legal question - whether the original unauthorised training on copyrighted material that has already occurred constitutes fair use or infringement - remains contested across multiple ongoing cases in different US courts, with rulings so far producing an inconsistent picture that neither AI labs nor rights holders can treat as fully settled precedent, leaving significant legal uncertainty hanging over the training-data practices of even the companies that have settled specific claims. News publishers have pursued a more commercially pragmatic path than book authors in several cases, striking direct licensing deals with OpenAI, Google and other labs that provide the AI companies with clean, structured access to current content in exchange for disclosed and undisclosed payments, while book publishers and authors' groups have more often pursued litigation first and licensing negotiations second, reflecting different assessments of leverage and different economic stakes between an industry built on subscription and advertising revenue versus one built on unit sales of individual works. Indian publishers and authors, operating under a copyright regime that has not yet produced comparable high-profile AI-training litigation, have watched the US settlements as a preview of arguments that may eventually surface in Indian courts, particularly as Indian-language content becomes increasingly valuable training data for the Indic-language foundation models that Sarvam, Krutrim and international labs are all racing to improve. What to watch: whether any pending US fair-use ruling produces a decisive precedent that reshapes settlement dynamics for the remaining unresolved cases, whether Indian publishers or authors' associations initiate comparable legal action, and whether the emerging licensing-price benchmarks from settled cases become the basis for a broader industry-standard training-data marketplace.
Original source: Reuters