Anna’s Archive Calls on Volunteers to Digitally Preserve Books Before AI Companies Destroy Them

Open digital library initiatives are launching urgent volunteer campaigns to scan and preserve physical books before rare printed works disappear from public access. The push comes in response to a growing trend where artificial intelligence companies purchase physical literature in bulk, slice the bindings, and destroy the paper copies after rapid scanning. Activists behind projects like Anna’s Archive warn that without open preservation efforts, vast amounts of human knowledge risk becoming permanently locked inside proprietary corporate databases.
Why Volunteers Must Preserve Physical Books Now
The campaign highlights a stark shift in how text is digitized for machine learning. AI firms require massive datasets to train large language models. However, instead of sharing their digitized archives with the public, these corporations retain the scanned texts on private servers. As physical copies are bought out from second-hand markets and destroyed during high-speed ingestion, public access to physical books shrinks.
Volunteers behind open archiving efforts contend that time is running out. When an out-of-print book is bought, destroyed, and ingested exclusively into a closed model, the original work becomes practically inaccessible to independent researchers, students, and readers worldwide.
Destructive Digitization vs Non-Destructive Scanning
To understand why physical literature is disappearing, one must look at the technical mechanics of data ingestion. Traditional archiving relies on non-destructive scanning, where operators carefully turn pages beneath overhead cameras. While this method preserves the original physical book, it is time-consuming and expensive.
By contrast, high-speed industrial scanning relies on destructive page-severing:
- Spine Removal: Automated guillotines slice off the book’s binding, turning bound pages into loose sheets.
- Sheet-Fed Ingestion: Loose pages pass through high-speed double-sided document scanners in seconds.
- Automated Processing: Optical character recognition software processes raw images into clean text blocks for training datasets.
- Physical Disposal: Sliced paper is recycled or discarded immediately after scanning.
This industrial approach allows technology companies to digitize thousands of pages per hour. However, it permanently eliminates the physical source material while keeping the output restricted to internal corporate use.
Open Archives and the Threat of Knowledge Monopolies
Open digital library projects rely on distributed volunteer networks to counteract corporate data hoarding. Volunteers use overhead book scanners, flatbed devices, and custom camera setups to digitize literature without ruining the physical object. Once processed, these files enter open digital repositories, ensuring that readers and researchers worldwide maintain free access to the text.
Advocates argue that when commercial developers monopolize written works without public release, society suffers a net loss of access. If an out-of-print title is destroyed during corporate scanning and the resulting dataset remains proprietary, the work effectively vanishes from the public domain. Consequently, open archive advocates stress that crowdsourced scanning remains a critical line of defense against corporate information monopolies.
Explore more AI dataset training practices from 90Network.




