31
The Vanishing Page: AI Firms Scan Then Destroy Rare Book Editions
(dallasexpress.com)
Posts from the RSS Feed of HackerNews.
The feed sometimes contains ads and posts that have been removed by the mod team at HN.
Every fucking website has to find a scandalized take about the several companies buying one (1) of every book.
Most of them were equally scandalized when the same companies did the obvious thing and grabbed a shadow library torrent.
Almost none of them recognize that when they demanded things be done properly - that's what this is.
This is not proper either
Elaborate.
I don't recall anybody demanding AI training involve destruction of physical media
They can't resell it, or their scan isn't a "backup copy," it's just piracy again. Same reason they can't share their scans of apparently-rare books nobody's cared enough to put on Archive.org.
Cutting the spine off is the simple way of scanning a mass-produced book. The alternatives are Rube Goldberg devices, or abundant tedious human labor.
What are you asking for, if not this? A private library of obscure forgotten works? A warehouse full of carefully-preserved dead weight? This was always the immediate alternative, and we told you as much, when y'all huffed about simple piracy. Like the average person on Lemmy gives two shits about piracy, outside this context.
I asking for them to go out of business if their business model is destroying art to manufacture slop
When the bubble bursts, this tech isn't going away.
Last year some blog recreated GPT-2, the first LLM 'too dangerous to release!,' for twenty bucks. Corpus to model within the hour. The billions spent on research since then have let newer models of similar size benchmark a hair below the big-iron cutting edge.
I recently found out Chroma, once a leading image model, was made by like two people, for presumably $200,000. Not nothing... but within reach of even small organizations. Recent pressure has been toward making models even smaller and less structured than that. Several updated video and editing models have quietly declined to release their weights, because they're small enough and powerful enough that most people would just run them locally. That desire, and that potential, will not vanish for wishful thinking.
If models had to train on nothing but public-domain, BSD, and CC0 materials, they'd still work the same way. If the money dries up to-morrow, then what's possible with hobbyist research alone will continue. If even they hit a brick wall, barely past current capabilities, that's still a couple gigs of linear algebra that'll try and do anything you ask.
Meanwhile - people will still make art. People make art even when it's illegal. Photography didn't stop anyone from drawing.