you are viewing a single comment's thread
view the rest of the comments
[–] 7 points 2 days ago (5 children)

It was posted like some kind of gotcha but....this book was in the dataset.

I'm aware that LLMs don't keep the dataset in memory but it "knows" that this is an existing work but these sites aren't doing any wild calculations to figure out if it was AI-generated. They basically just check to see if the sentences exist elsewhere and they do. In the original dataset.

So it was wrong to say it's AI-generated but it correctly identified that it's not original.

  • source
  • parent
  • hideshow 10 child comments
  • [–] 94 points 2 days ago (1 child)

    They analyze statistical patterns, they don't cross reference the training data.

    It's picking up the text as generated because it's mostly guess work and constantly spits out false positives, especially with non native speakers.

  • source
  • parent
  • hideshow 2 child comments
  • [+] -7 points 2 days ago (2 children)

    Honestly I didn’t expect them to flag non-native speakers. LLMs don’t often make gramatical mistakes (or at least not ones your average joe could identify)

  • source
  • parent
  • hideshow 4 child comments
  • [–] 50 points 2 days ago (3 children)

    Foreign speakers often know English better than most Americans

  • source
  • parent
  • hideshow 6 child comments
  • [–] 7 points 2 days ago* (1 child)

    Oooh yeah, the most common dumb mistakes are (by my own checking) done by Americans in the VAST majority cases. Stuff like not knowing how to use your/you're, there/their/they're, cloths/clothes, definitely/defiantly, where/were/we're, its/it's etc. correctly and the many other failings of basic grammar like "would/could of" and so much else.
    I don't think there's another country in the world where the native speakers are so fucking bad at even the most basic shit.
    The most common response I've gotten and seen when others correct stuff like that is "calm down, English probability isn't their main language" yet when checking it's been Americans who made the mistakes 99% of the time.

  • source
  • parent
  • hideshow 2 child comments
  • [–] 5 points 2 days ago* (1 child)

    I was going to say:

    The difference is between growing up, learning a language by hearing and making the sounds, and learning it as a foreign language, having to understand and memorize the rules.
    "They're, their, there" all sound the same as well as "could've, could of".

    But interestingly other English speakers don't make those mistakes as often as people emerging from the american education system.

  • source
  • parent
  • hideshow 2 child comments
  • [–] 2 points 1 day ago (1 child)

    It's plausible that Americans are worse at this than foreigners, but English spelling is a crime and everyone who makes mistakes ought be excused on account of how ridiculous the "rules" are.
    I know how to spell most of the words I use and I resent the fact that my memory is polluted with this otherwise useless knowledge.
    In many other languages if you know how to say a word, you automatically know how to spell it. We need a world language with grammar as simple as English and spelling as straightforward as German.

  • source
  • parent
  • hideshow 2 child comments
  • [–] 1 point 1 day ago (1 child)

    Come learn french if you think english spelling is convoluted. And doesn't the language you're looking for exist ? Doesn't esperanto check all your requirements ?

  • source
  • parent
  • hideshow 2 child comments
  • [–] 1 point 1 day ago

    I almost mentioned French, but thought it didn't need to catch strays here haha.
    To be honest I don't know much about Esperanto, but it lacks one critical requirement - to be a world language. And realistically it has no chance to be. The number of people who speak it is a rounding error.

    I think what has a real shot at happening instead is that English spelling simplifies gradually and organically until it is largely straightforward. That is if Chinese doesn't end up as the next world language lol

  • source
  • parent
  • [–] 3 points 2 days ago (1 child)

    The decline I've seen in my time on the internet is genuinely shocking

  • source
  • parent
  • hideshow 2 child comments
  • [–] 32 points 2 days ago (1 child)

    That is not in the least bit how a tool like this works.

    All AI detection is extremely unreliable, but they operate on principles which, if the assumptions supporting them were true, would be sound. The way you imagine they work is different: you're describing a "novel text detector" which is not at all the same thing as an "AI text detector".

    AI detection works, at a very high level, by throwing a bunch of examples of AI text, and a bunch of examples of non-AI text, into a machine-learning model, and training it until it is able to recognise the AI examples as such. It doesn't work by throwing in all existing text including novels written before AI, because that would never do what you want.

    (The reason, if you're interested, why this ends up not working is because the features such a model identifies as indicative of AI text are very sensitive: if you train it on Claude and ChatGPT, it will not work on Gemini output. If you train it on Gemini, when Gemini updates it will get worse. If someone generates text with a weird prompt, it may slip by. If someone writes in a weird way, it may get flagged. And if any AI company wanted to defeat AI detectors, it could trivially feed its output through one during the training process and give that output as examples to avoid in the training, so that the model learns to create output which doesn't "look like" AI output to those detectors.)

  • source
  • parent
  • hideshow 2 child comments
  • [–] 2 points 2 days ago (1 child)

    I was being overly simplistic - I meant more that the patterns it's trying to detect were created by an LLM trained on the data they're inputting.

    It's like how reddit comments from 2016 look generated. If you stick one of those into an AI detector, it'll give a false positive for the same reason

  • source
  • parent
  • hideshow 2 child comments
  • [–] 4 points 2 days ago (1 child)

    I'm sorry, what? It's correctly identified that it's what, now?

  • source
  • parent
  • hideshow 2 child comments
  • [–] -2 points 2 days ago (2 children)

    It's original in the book. What the user entered is an exact copy.

    "Original" is not the correct work but you know what I mean.

  • source
  • parent
  • hideshow 4 child comments
  • [–] 5 points 2 days ago (1 child)

    The little scale thing doesn't say it is 100% not original, it says it is AI/LLM generated.

  • source
  • parent
  • hideshow 2 child comments
  • [–] -2 points 2 days ago (2 children)

    Yes...because the pasted text has some of the exact patterns that are part of the dataset that the website trained on...because that dataset contains text generated from a dataset containing the exact paragraph...

    I feel like you're being deliberately obtuse.

  • source
  • parent
  • hideshow 4 child comments
  • [–] 6 points 2 days ago* (1 child)

    I feel like you don't understand the difference between 'AI generated' and 'something that existed over 200 years ago'.

  • source
  • parent
  • hideshow 2 child comments
  • [–] -3 points 2 days ago (1 child)

    I feel like you don't understand that LLMs train on text that existed 200 years ago...

  • source
  • parent
  • hideshow 2 child comments
  • [–] 1 point 2 days ago (1 child)

    So you confirm that you don't understand the difference. Ok.

  • source
  • parent
  • hideshow 2 child comments
  • [–] -2 points 2 days ago (1 child)

    Sorry for not answering your idiotic question. Also sorry that you're dead set on deliberately misunderstanding then doubling down on it. Hope it makes you feels smart!

  • source
  • parent
  • hideshow 2 child comments
  • [–] 0 points 2 days ago (1 child)

    Everything you said was hallucinated. Nothing you said was even slightly coherent or connected with reality. Absolutely nothing was correct. If you let your cat dance on the keyboard more sense would come out.

  • source
  • parent
  • hideshow 2 child comments
  • [–] 3 points 1 day ago (1 child)

    That must have sounded really clever and powerful in your head for you to hit send, huh? Couldn't decide which zinger to go with so you sent all 4?

    I'd have given you more credit if you just posted the quote from Billy Madison. Try harder lol

  • source
  • parent
  • hideshow 2 child comments