Legacy Media vs. AI: The Licensing Deal Rush
The relationship between journalism and artificial intelligence has shifted rapidly from skepticism to a high-stakes commercial sprint. As generative AI models like ChatGPT require massive amounts of high-quality text to learn and improve, tech giants are now paying for the privilege of accessing it. For legacy media companies, this represents a pivotal moment. They must decide whether to sue AI companies for copyright infringement or sign lucrative licensing agreements to secure a new revenue stream.
The billion-dollar scramble for high-quality data
For years, search engines and AI crawlers scraped the open web with little resistance. However, the launch of ChatGPT changed the equation. Publishers realized their archives were being used to train products that could eventually replace them. In response, OpenAI and other tech leaders have started writing checks.
The goal for OpenAI is clear: secure legal access to reliable, fact-checked, and human-written content to reduce “hallucinations” (errors) in their models. For publishers, the goal is survival and monetization.
Major players signing on the dotted line
The first half of 2024 saw a flurry of high-profile partnerships. These deals grant AI companies permission to ingest current reporting and decades of archives.
- News Corp: In May 2024, News Corp (owner of The Wall Street Journal, New York Post, and The Times) announced a landmark multi-year partnership with OpenAI. Reports estimate the deal could be worth over $250 million over five years. This gives OpenAI access to premium content that is otherwise locked behind strict paywalls.
- Axel Springer: As one of the first major movers in December 2023, the German publishing giant licensed content from Politico and Business Insider to OpenAI. This deal was significant because it set the precedent that AI companies were willing to pay for current news, not just historical data.
- Dotdash Meredith: The publisher behind huge lifestyle brands like People, Better Homes & Gardens, and Investopedia signed a deal in May 2024. This partnership focuses on enhancing OpenAI’s models with specific lifestyle and intent-based data, which is different from hard news.
- Vox Media and The Atlantic: Both organizations recently joined the ranks, allowing their articles to be surfaced within ChatGPT with attribution.
The holdouts: The New York Times vs. OpenAI
Not everyone is rushing to sign a contract. The New York Times chose a different path by filing a lawsuit against OpenAI and Microsoft in late 2023.
The Times alleges that millions of its articles were used to train chatbots that now compete with it as a source of reliable information. Their legal complaint included examples where ChatGPT purportedly regurgitated large chunks of paywalled articles verbatim.
This legal battle highlights the core tension in the industry. The Times argues that using their work without permission is copyright infringement, regardless of whether a deal is offered later. If The New York Times wins, it could force AI companies to delete data or pay massive damages. However, if the courts rule that training AI falls under “fair use,” publishers effectively lose their leverage to demand payment.
What do publishers actually get?
The structure of these licensing deals usually involves two main components: cash and traffic.
- Upfront and Annual Payments: The financial terms are often the headline. For struggling media companies, an injection of $50 million or $100 million over a few years is a lifeline. It provides a stable revenue source that isn’t dependent on the volatile advertising market.
- Attribution and Links: A major fear for publishers is the “zero-click” future, where a user asks a chatbot a question and gets an answer without ever visiting the source website. To combat this, deals like the one with The Atlantic require OpenAI to provide clear citations and links back to the original article when that data is used to generate an answer.
- Technology Access: Many of these agreements are reciprocal. In exchange for data, newsrooms get early access to OpenAI’s enterprise technology. This allows them to experiment with AI for tasks like headline testing, summarizing archives, or data journalism.
The Reddit and Google factor
It is not just OpenAI making moves. Google is also aggressively securing data for its Gemini models and “AI Overviews” in search results.
In early 2024, Google struck a deal with Reddit reportedly worth $60 million annually. This gives Google real-time access to Reddit’s “firehose” of user-generated conversations. For Google, Reddit provides the colloquial, human-verified answers that searchers often crave. For Reddit, it monetizes their user base in a way that advertising alone could not.
The risk of a "Walled Garden" internet
While these deals are good for the bottom line of large media conglomerates, they raise questions about the future of the open web.
If AI models are primarily trained on content from partners who have signed exclusive deals, smaller publishers and independent bloggers may be left behind. Their content might be excluded from the “verified” training sets, meaning they receive less visibility in AI-generated answers.
Furthermore, as publishers lock down their sites to prevent unauthorized scraping (using tools like robots.txt or tougher paywalls), the internet becomes more fragmented. Information that was once searchable via a simple Google query might soon reside inside the “black box” of a chatbot, accessible only because a corporation paid a licensing fee behind the scenes.
Frequently Asked Questions
Why are publishers suing AI companies? Publishers like The New York Times argue that AI companies used their copyrighted articles to train models without permission or payment. They claim this creates a product that competes with them using their own stolen labor.
How much are these licensing deals worth? The values vary significantly based on the size of the publisher. The News Corp deal is reportedly worth over $250 million over five years, while Reddit’s deal with Google is valued at roughly $60 million per year.
Will ChatGPT link back to the news articles? Yes, most of the new licensing agreements specifically include provisions for attribution. When ChatGPT uses information from a partner like The Atlantic or The Wall Street Journal, it is designed to provide a link to the original source.
What happens if I don’t sign a deal? Publishers who do not sign deals face a difficult choice. They can block AI crawlers to protect their content, but this means they might be excluded from AI-generated answers entirely. Alternatively, they can leave their sites open and risk having their content used for free under “fair use” defenses, unless current lawsuits change the legal standards.