Millions for a comment. Why are IT giants buying up books and corporate emails?

In the era of rapid generative AI development, human data has become a critical resource for training neural networks. The article analyzes the current market situation where IT giants like OpenAI are investing heavily to gain access to high-quality data sets. This process includes not only paid APIs from social platforms like Reddit, but also multi-million dollar deals with media conglomerates, as well as the mass acquisition of copyrights for printed books and corporate email archives. The authors emphasize that information has become the 'new oil' for tech companies striving to improve content generation quality. This strategy raises questions about copyright, privacy, and the ethics of using intellectual property to train models. Ultimately, the battle for data is becoming a defining factor in AI industry competition, where access to unique content directly impacts the efficiency and market value of the systems being developed.
This is a summary. Read the full article at the original source:
HabrRelated stories
From a Single Agent to an AI Holding: Scaling Autonomous AI Systems
This article explores the evolution of autonomous AI systems, moving from executing simple individual tasks to building comprehensive AI organizations…
Splice CEO Kakul Srivastava warns that AI-generated emails are eroding authentic communication
Kakul Srivastava, CEO of the music production platform Splice, has expressed concerns regarding the impact of generative AI on human interaction. In a…
David Robinson, a former safety researcher at OpenAI responsible for drafting safety reports for major model releases, has officially resigned from th…



