Executives working on AI at Microsoft and OpenAI admitted what its critics have been saying all along: Large language models are predatory pieces of technology that have been built on what a Microsoft executive called “an astonishing theft of unprecedented proportions,” and the “largest theft of labor in human history.” An internal Microsoft document said generative AI products have created a “doom loop” that is killing “the entire web.”

Those statements and a series of other mask-off moments feature heavily in an unredacted court filing that was unsealed Thursday in the behemoth New York Times vs OpenAI copyright lawsuit that has been winding its way through the court system for years. In a filing asking for summary judgment (basically, a filing with the court asking it to rule), lawyers for the New York Times laid out a series of admissions made by Microsoft and OpenAI executives in documents and depositions that until now had remained either sealed or redacted at the request of Microsoft and OpenAI.

It’s easy to see why the AI companies wanted to hide this from the public. The statements, taken together, are some of the most damning indictments of the ways LLMs were trained, how they worked, and the immediate threat they pose to human labor. It is a reminder that even as AI becomes more powerful and companies try to shift the narrative to the supposed existential risk of “superintelligent” AI, the tools they have already built were created by stealing from human creativity and labor and are by definition existential threats to the human labor market.


you are viewing a single comment's thread
view the rest of the comments
[–] -2 points 12 hours ago (8 children)

a “doom loop” that is killing “the entire web.”

the corporate web maybe. how does it affect independent parts like the fediverse?

  • source
  • hideshow 8 child comments
  • Something others haven't touched on is the burden all the scrapers are putting on small websites and forums.

    Lots of niche/hobbyist site runners have repeatedly talked about how the scrapers inefficiently scrape the same pages, images, and files endlessly. Until they crash the site, run them out of bandwidth, or run up their bills till they cant afford to host a site they've been running for many years before this nonsense.

    All in the futile pursuit of more to dump into the machine because of the belief if they just had more data all the issues with the models would finally be solved.

  • source
  • parent
  • [–] 14 points 12 hours ago (4 children)

    By bombarding it with bots like Reddit? Honestly, I can't tell what is bot and not. Are you a bot? Am I a bit? Who knows?

  • source
  • parent
  • hideshow 4 child comments
  • [–] 7 points 12 hours ago (2 children)

    Yeah once the mainstream is fully saturated they'll turn their gaze to smaller, more "authentic" places. Imagine if they could have a boot infiltrate a tiny little private chat between a dozen people by building up the appearance of an organic presence on the fediverse / indie web.

    Now advertisers (and potentially, malicious state actors) have a little spy and influencer in your group.

  • source
  • parent
  • hideshow 2 child comments
  • [–] 6 points 12 hours ago

    Slop commenting. The litany of AI slop websites that have popped up, polluting search results. The absolute flood of AI slop pull requests being submitted to open source projects, to the point at which many open source maintainers are buckling under the weight of garbage being thrown their way. Bots everywhere. The list goes on and on and on.

    The internet is already a dessicated husk of what it once was.

  • source
  • parent