Amazon is facing fresh scrutiny over the use of physical books for artificial intelligence training after an investigation reportedly tracked a shipment of hard-to-find titles to a company facility in Las Vegas, where workers allegedly cut the bindings from books before scanning their pages and eventually destroying them.
The investigation, reported by 404 Media, involved placing a tracking device inside a book from a large order and following its journey across the United States. The shipment ultimately reached an Amazon facility associated with an operation known as VGT3.
The discovery has raised concerns among booksellers, librarians and readers about what happens to physical books once they are acquired as potential sources of AI training data.
It’s because that might allow them to monopolize access to information later on. Your future library is called Amazon and it costs you a 100 bucks per month subscription. Plus, they can change the information displayed to you at will, as it all comes from their algorithmic black box.
First off, if my future Library of Everything is called anything other than “Alexandria”, I’ll be quite cross
Sadly, they might just call it Alexa-Libria or some sht
They aren’t doing this to gatekeep information, they’re doing this to rewrite history. Fascism cannot survive otherwise.
Words cannot express how much I fucking hate corporate AI
I’m flirting with radicalization over here
Was global warming not enough?
It should have been
Welcome to the movement. This is a good book to get started. I would recommend you to skip straight to chapter 1.
Thank you!
You are welcome :>
Fascism*

Modern day book burningHow is this in any way comparable to that? The bad thing abt nazi book burning wasnt the destruction of physical copies of books, but the destruction of knowledge and their intent behind it
hard-to-find titles
but the destruction of knowledge
I think their point is that scanning preserves the content rather than destroying it, but I’m guessing the scans aren’t exactly public.
Destroying the original is flat out stupid from every angle unless it’s unavoidable or destruction (or taking control) is the goal.
Honestly it might be worse if that’s not the idea, but yeah it’s still nazi coded at best
The same applies here
If these LLMs are so great why do they need to destroy the books to scan them in.
My phone can scan a page and account for the curve caused by the spine but it’s too much for their LLM data centers to figure out.
They’re feeding it in, like a copier machine – one sheet at a time so they lop off the binder, run the stack, and send it to the trash or furnace or wherever.
But yes, they could be doing it non-destructively, but they’re being fascist-like.
It’s faster, cheaper and can be automated. That’s it. That’s the whole reason. To avoid paying a person to do it.
Because if you get all the copies of a rare book, scan it and destroy it, then nobody else can use it for training but you. They don’t have to, they choose to.
Most “rare” books exist in digital format.
Then they are not rare by definition. These are not the books they are buying.
Yeah they are. The company needs to scan their own physical copy, for legal reasons.
They did it your way, and people screamed bloody murder about “piracy,” so now they’re doing this. Anyone who demanded they do things “the right way” meant, this. Whether they knew it or not.
the first person that actually speaks sense in this thread
At the core of it is a rotten person making rotten decisions. That’s why each solution is worse than the previous.
According to David Gerard (youtube podcast link) they cut off the spines of the books to allow for rapid scanning
Edit: Ars Technica too, thanks alapakala
Yes, but being the only one able to get their hands on that book and destroying it has that added benefit. My point is that there is no motivation to preserve those books, quite the opposite in fact.
My phone can scan a page
Great, now do that a hundred more times, as quickly as possible… for every book in an entire palette of books. This week.
And then you can’t ever sell the books, or give them away, or lend them out, because then your scans are not “backup copies.” It’s just piracy.
I have been explaining this to people for an entire year, and this stupid fucking story still gets rehashed every week.
If you destroy the book ya don’t fucken own it, nimrod.
Yes… you do. That’s why it’s a backup copy.
If you think copyright law shouldn’t work this way, great. We should should really change it to remove any incentive to hoard these scans - obviously these purchases aren’t making the authors any more money. But right this second, this is the law, and that’s what following it looks like.
Try that argument with a dvd, or a music cd, or pretty much anything and see how fast you get busted for “piracy”
You mean like ripping a CD to MP3s? Which is obviously legal?
Or backing up a DVD to VHS, which is supposed to be explicitly protected, and created a legal argument for DeCSS defeating anti-copying measures?
And I’ve also been explaining to people that destroying the book means losing the rights to the contents of said book for an entire year. Yet, here we are
You have been lying to people.
Please explain
If you have a backup - and the original gets destroyed - you still own the backup.
That’s what makes it… a backup.
Yes, the backup still exists. What is also true is that you gave your original away or destroyed it intentionally, as it is in this case, so you lose the right to the backup.
Giving it away means there’s two copies owned by different people. That’s why you have to delete your backups.
Destroying it, accidentally or otherwise, means there’s still one copy owned by whoever bought it. That’s why you get to keep your backups. What the fuck else do you think a backup is?
A backup is not treated as meaningfully different from the original copy. That’s why you’re allowed to make one, at all. That’s also why you can keep the backup - you still own that copy. Even if you’ve transferred formats. Which is also explicitly allowed.
Why cant they just warehouse the books until the copyright expires, then? Why don’t they go slower? Because they don’t want to and don’t care about the damage.
Don’t stick up for them, they aren’t gunna pay you for that year you spent trying to hold water for them. Nobody is cuz we don’t want that either way.
Yeah why don’t they just spend an order of magnitude more effort and money to maintain an unusable forest of dead trees until *checks notes* forever minus a day?
I fucking hate bad arguments, for my own sake. I hate misinformation going round and round, with shrill comparisons to the fucking Nazis, because someone dared to buy and privately preserve a copy of a book that nobody else was likely to ever purchase.
Relax copyright if you want this to go better. The idiots inflating the bubble are doing this because our stupid laws say they kinda have to. It’s an asinine legal situation that leads to crap like this, and no amount of well-they-oughtta is gonna change crap like this.
I feel like we’re missing information from that summary:
- why are they “hard to find” titles?
- are the books being sold digitally under a license?
- is that license for 1 digital copy, or is the same book being copy-pasted to each buyer?
- how were the books acquired?
- do the books have redistribution rights?
They never list titles or condition of the books so I just assume they are the same books you can find in big piles at the dump mostly old text books and technical books. I love books I own hundreds of them, but people don’t buy enough 2nd hand books there are sites where you can buy books by the yard so you can have a neat and matching library, there are warehouses full of books to rent for movie sets. Noone realy cares they just want to be upset.
They are buying physically rare books from antiquarians with online listings. Those books are all out of print, or they’d just buy a new copy that’s readily available and surely cheaper. Rights are irrelevant because the entire point is to create a digital backup copy… in accordance with copyright law.
It’d be great if they could toss these scans onto the Internet Archive. But they can’t. Legally that’s the same as piracy, which is the whole reason they’re doing this, instead of just mirroring Anna’s Archive. That was their first whack and people were aghast. Funny how they’re not any happier, now that these companies are doing things “the right way.”
I do think piracy is better than destroying antiques. They could both pay for the books they copy and not destroy rare ones that might be unique. There’s no contradiction here.
Realistically that’s not going to happen unless they can pass on the physical books and/or their scans. They might bother paying for a basement-full-of-crates “library” if they got to put the PDFs online and brag about preservation. Same deal for dumping truckloads of obscure nonsense onto local libraries and loudly scoffing at critics of that let’s-say-charity.
But under the rules they have to follow, what they’re doing is obviously what they were always going to be doing. They only want one copy of each book ever mass-produced. They’re scanning them in the easy and fast way that also solves any legal concerns. They kept it quiet because people predictably flip out over this wide-but-shallow purchase.
Their actions are unlikely to change unless the legal situation changes. Let’s maybe change it for the better, for all of us.
This has nothing to do with privacy. It’s published works from decades ago.
All of these texts should be public-domain by now.
If anything here is eyebrow-raising, it’s the store that slipped a tracking tag into a shipment of books to unmask a straw buyer.






