AI工具Score B (60)
AI companies are destroying books. What will they come for next? - The Globe and Mail
2 小时前2 viewsSource: theglobeandmail.com
Open this photo in gallery: Books line shelves at the North York Central Library in Toronto in February, 2024. Chris Young/The Canadian Press Share Save for later Please log in to bookmark this story. Log In Create Free Account Artificial intelligence is not only changing the world as we know it, but it’s destroying parts of our world along the way. The latest casualty? Books. Anthropic, to train its Claude AI model, has been buying physical books in order to scan them and ingest their contents. To do this as efficiently – and legally – as possible, special machinery cuts the pages from their bindings and scans them. The books’ remnants are then discarded. According to reports citing court documents , Anthropic has purchased – and destroyed – millions of physical books in the process. By acquiring physical books, the company owns them – and thus does not have to worry about copyright infringement the way it would by downloading pirated libraries, also known as shadow libraries. In a landmark class-action copyright lawsuit, Bartz v Anthropic, authors (including Andrea Bartz) argued that books were being used without permission to train LLMs, accusing the company of downloading millions of pirated books. The company settled after the judge’s ruling last year. This month, the US$1.5-billion settlement was given final approval by a judge in San Francisco, sending the online discourse into a fever pitch. Amazon rolls back most flagship AI models in strategy overhaul, Business Insider reports Still, last year’s court decision was seen as a big win for the AI industry. Because a (now retired) judge, William Alsup, ruled that AI training was “quintessentially transformative” – that the AI models were trained not to replicate the books, but to create something new. And that using books without permission to train AI was fair use if they were acquired legally . So AI companies can train their algorithms using copyrighted works, as long as the company pays for the books – not by licensing them, but by buying individual copies. As per U.S. copyright law , the first sale doctrine states that once a copyright holder has authorized the sale of a physical copy, their control over that specific object ends. The purchaser has “the right to sell, display or otherwise dispose of that particular copy , notwithstanding the interests of the copyright owner.” Perhaps ironically, the destruction is necessary in order for the use to be legal. So companies were given the okay to buy individual books and scan their contents with impunity. And they have been busy shopping, copying and dumping. According to the Washington Post , Anthropic has purchased millions of books in pursuit of their content, often in batches of tens of thousands. The Post cites a document describing a “hydraulic powered cutting machine” that would “neatly cut” books, whose pages would be scanned, and ultimately the scanning company would “schedule with the recycling company to pick up the completed books.” Opinion: China is AI-maxxing, and it has a lot to teach us This month, 404 Media reported that some online booksellers have seen a “historic” surge in sales this year. And that ISBNdb, a book database that offers high-volume acquisition services, is helping AI labs with bulk purchases of between 1,000 and 1 million (!) books per order. An article on the ISBNdb website titled “The Receipt is the new License” called Anthropic’s method “an escape hatch” – “so unusual, so methodical, and so deliberately low-profile that it reads more like a Cold War logistics program.” Indeed, the Post reported that it was code-named “Project Panama.” It’s not just Anthropic; sourcing books is a widely used approach for training LLMs across the AI industry. To be fair, we’re not talking about the Book of Kells here. These books are available on the commercial market. Still, they may be rare. What is lost when a conglomerate destroys the final copy of, say, a little-known biography, a collection of obscure fairy tales, or the history of a small town that has since been swallowed up by urban sprawl? This knowledge is then lost to history, devoured by Claude, Meta, and the like. ( Elon Musk says he has instructed his AI team not to destroy any books it uses in such a process. Hero.) For authors, even those who will make US$3,000 or so in the Bartz settlement, this is maddening. It takes years to write a book, and to have it destroyed in seconds so companies can make gazillions off the backs – not to mention the blood, sweat and tears – of the people who did the actual research and wrote and edited the actual words is galling and unjust. At the very least, the AI companies should be paying to license these books. Worse, it feels like some sort of metaphor – or portent. As the German poet Heinrich Heine wrote , prophetically, in 1820: “When they burn books, they will, in the end, burn human beings too.” So what happens when they destroy books? AI is here to stay. Some of the literary history it burns through to feed it may not be as fortunate.
Read the full original article:
theglobeandmail.com