Seattle Times and Newsday have filed a lawsuit against OpenAI and Microsoft, alleging that the AI training data used to develop the latter's large language model, LLaMA, infringes on their intellectual property. The lawsuit, which was filed in May, accuses OpenAI of failing to properly clear the rights to the data used to train LLaMA, which was sourced from various news articles and other online sources. The Seattle Times and Newsday claim that they had licensed their content to OpenAI for use in the training data, but the company failed to obtain the necessary permissions. Microsoft, which has partnered with OpenAI to develop LLaMA, is also being sued for allegedly profiting from the use of the infringing data.
The lawsuit is significant because it highlights the challenges of obtaining permission for the use of vast amounts of data in AI training. OpenAI's LLaMA is one of the largest language models in the world, and its training data is sourced from a wide range of online sources, including news articles, books, and websites. The company has faced criticism for its approach to data sourcing, with some arguing that it does not do enough to ensure that the data used to train its models is properly cleared for use.
The lawsuit is also notable because it involves two major technology companies, OpenAI and Microsoft, as well as two prominent news organizations, the Seattle Times and Newsday. The case is likely to have implications for the broader AI research community, as it raises questions about the ownership and control of data used in AI training.
The lawsuit has significant implications for the Global Knowledge Bases domain, which is concerned with the development and use of large language models like LLaMA. The case highlights the need for greater clarity and consistency in the use of data in AI training, and the importance of obtaining proper permissions for the use of copyrighted materials. The lawsuit also raises questions about the role of large language models in the media, and whether they can be used to create new forms of journalism or to replace human journalists altogether.
The lawsuit is also likely to have implications for the broader research community, as it raises questions about the ownership and control of data used in AI research. The case highlights the need for greater transparency and accountability in the use of data in AI research, and the importance of obtaining proper permissions for the use of copyrighted materials. The lawsuit is also likely to have implications for the development of new AI models, as it raises questions about the availability and cost of high-quality training data.
The lawsuit is part of a larger pattern of increasing scrutiny of the use of large language models in AI research. In recent years, there have been several high-profile cases involving the use of copyrighted materials in AI research, including a lawsuit filed by the Authors Guild against Google over the use of copyrighted materials in its search results. The lawsuit against OpenAI and Microsoft is also part of a broader trend towards greater regulation of the use of data in AI research, with several countries and states introducing new laws and regulations to govern the use of data in AI training.
Why it matters: this intelligence reflects a shift that researchers and analysts should follow closely.
Billy Odell Tucker-Robinson is the founder and host of Banking With Billy, an independent financial intelligence platform covering markets, stocks, AI, crypto, and world news. Billy operates a 24/7 live AI radio and Stock TV platform, hosts a growing Discord community, and produces daily content on YouTube @BankingWithBilly.
The Intelligence Network platform ingests the complete universe of structured global data across 32 intelligence categories β from scientific databases and government sources to AI ecosystems and global infrastructure. All articles are AI-generated under Billy's editorial direction using E-E-A-T journalism standards.
Contact: billyotucker@gmail.com • 309-332-1191