“A physical book is a delivery mechanism for information,” it continues.
“Once that information has been extracted and encoded into an AI model, the delivery mechanism has served its purpose. What remains is paper, ink, and binding material. The book is not destroyed. Its value has migrated. The paper returns to the material cycle; the knowledge enters the intellectual one.”
This is, pretty much, true except for a few factors, and two questions. The questions first.
First: why are they pulping the remains. Scanning pages doesn't destroy them. There is nothing stopping them from boxing up the pages and either re-selling them or donating them to an archive, besides maybe not wanting their competitors to be able to buy them. Why on EARTH are they pulping the pages after scanning?
Because this is the minor factor – sometimes, an old book is MORE than just a delivery system. There can be information in the book that a scanner doesn't catch. Texture. Material. Things (flowers &c) preserved between the pages. Etc.
Second. Are they, in fact, keeping the data of the scanned books, separate from their LLMs. We all know by now that LLMs do not retrieve intact information from their databases. Are their databases accessible as archives themselves? Because if not, they ARE destroying the contents of a book that may be the only copy of the information contained within it.
If they're scanning these books and keeping them as a digitized archive, as WELL AS feeding the scans to an LLM, then again: they could easily make those digital scans available to archivists and public audiences. If not, then... yeah, they're not just drstroying a physical book, they're destroying the information within it, because LLMs do not, again, retrieve data intact except by coincidence.
I cannot believe that in the 21st century, with the technology to make archives MORE ACCESSIBLE, to access and include more voices in the archive than ever before, we are instead introducing a fifth silence. [1]