top 50 comments

sorted by: hot top controversial new old
[–] 205 points 2 years ago (27 children)
  • [–] 85 points 2 years ago (24 children)

    The results of this new GSM-Symbolic paper aren't completely new in the world of AI research. Other recent papers have similarly suggested that LLMs don't actually perform formal reasoning and instead mimic it with probabilistic pattern-matching of the closest similar data seen in their vast training sets.

    WTF kind of reporting is this, though? None of this is recent or new at all, like in the slightest. I am shit at math, but have a high level understanding of statistical modeling concepts mostly as of a decade ago, and even I knew this. I recall a stats PHD describing models as "stochastic parrots"; nothing more than probabilistic mimicry. It was obviously no different the instant LLM's came on the scene. If only tech journalists bothered to do a superficial amount of research, instead of being spoon fed spin from tech bros with a profit motive...

  • source
  • parent
  • hideshow 24 child comments
  • [–] 45 points 2 years ago (19 children)

    It's written as if they literally expected AI to be self reasoning and not just a mirror of the bullshit that is put into it.

  • source
  • parent
  • hideshow 19 child comments
  • [+] 39 points 2 years ago* (last edited 6 months ago) (18 children)
  • [–] 11 points 2 years ago (3 children)

    …a spellchecker on steroids.

    Ask literally any of the LLM chat bots out there still using any headless GPT instances from 2023 how many Rs there are in “strawberry,” and enjoy. 🍓

  • source
  • parent
  • hideshow 3 child comments
  • load more comments (2 replies)
  • load more comments (14 replies)
  • load more comments (3 replies)
  • [–] 94 points 2 years ago* (5 children)

    One time I exposed deep cracks in my calculator's ability to write words with upside down numbers. I only ever managed to write BOOBS and hELLhOLE.

    LLMs aren't reasoning. They can do some stuff okay, but they aren't thinking. Maybe if you had hundreds of them with unique training data all voting on proposals you could get something along the lines of a kind of recognition, but at that point you might as well just simulate cortical columns and try to do Jeff Hawkins' idea.

  • source
  • hideshow 5 child comments
  • [–] 45 points 2 years ago

    LLMs aren't reasoning. They can do some stuff okay, but they aren't thinking

    and the more people realize it, the better. which is why it's good that a research like that from a reputable company makes headlines.

  • source
  • parent
  • load more comments (4 replies)
    [–] 90 points 2 years ago (16 children)

    Did anyone believe they had the ability to reason?

  • source
  • hideshow 16 child comments
  • load more comments (13 replies)
    [–] 56 points 2 years ago (11 children)

    So do I every time I ask it a slightly complicated programming question

  • source
  • hideshow 11 child comments
  • [–] 19 points 2 years ago (10 children)

    And sometimes even really simple ones.

  • source
  • parent
  • hideshow 10 child comments
  • [–] 9 points 2 years ago (9 children)

    How many w's in "Howard likes strawberries" It would be awesome to know!

  • source
  • parent
  • hideshow 9 child comments
  • [–] 9 points 2 years ago* (4 children)

    So I keep seeing people reference this... And I found it curious of a concept that LLMs have problems with this. So I asked them... Several of them...

    Outside of this image... Codestral ( my default ) got it actually correct and didn't talk itself out of being correct... But that's no fun so I asked 5 others, at once.

    What's sad is that Dolphin Mixtral is a 26.44GB model...
    Gemma 2 is the 5.44GB variant
    Gemma 2B is the 1.63GB variant
    LLaVa Llama3 is the 5.55 GB variant
    Mistral is the 4.11GB Variant

    So I asked Codestral again because why not! And this time it talked itself out of being correct...

    Edit: fixed newline formatting.

  • source
  • parent
  • hideshow 4 child comments
  • load more comments (4 replies)
  • load more comments (4 replies)
  • [–] 56 points 2 years ago (4 children)

    The tested LLMs fared much worse, though, when the Apple researchers modified the GSM-Symbolic benchmark by adding "seemingly relevant but ultimately inconsequential statements" to the questions

    Good thing they're being trained on random posts and comments on the internet, which are known for being succinct and accurate.

  • source
  • hideshow 4 child comments
  • [–] 46 points 2 years ago (6 children)

    statistical engine suggesting words that sound like they'd probably be correct is bad at reasoning

    How can this be??

  • source
  • hideshow 6 child comments
  • [–] 19 points 2 years ago (2 children)

    I would say that if anything, LLMs are showing cracks in our way of reasoning.

  • source
  • parent
  • hideshow 2 child comments
  • [–] 45 points 2 years ago (3 children)

    They are large LANGUAGE models. It's no surprise that they can't solve those mathematical problems in the study. They are trained for text production. We already knew that they were no good in counting things.

  • source
  • hideshow 3 child comments
  • [–] 39 points 2 years ago (4 children)

    Are you telling me Apple hasn't seen through the grift and is approaching this with an open mind just to learn how full off bullshit most of the claims from the likes of Altman are? And now they're sharing their gruesome discoveries with everyone while they're unveiling them?

  • source
  • hideshow 4 child comments
  • [–] 51 points 2 years ago

    I would argue that Apple Intelligence™️ is evidence they never bought the grift. It's very focused on tailored models scoped to the specific tasks that AI does well; creative and non-critical tasks like assisting with text processing/transforming, image generation, photo manipulation.

    The Siri integrations seem more like they're using the LLM to stitch together the API's that were already exposed between apps (used by shortcuts, etc); each having internal logic and validation that's entirely programmed (and documented) by humans. They market it as a whole lot more, but they market every new product as some significant milestone for mankind ... even when it's a feature that other phones have had for years, but in an iPhone!

  • source
  • parent
  • load more comments (3 replies)
    [+] 36 points 2 years ago* (last edited 2 months ago) (2 children)
  • [–] 28 points 2 years ago

    This right here, this isn't conscientious analysis of tech and intellectual honesty or whatever, it's a calculated shot at it's competitors who are desperately trying to prevent the generative AI market house of cards from falling

  • source
  • parent
  • [–] 33 points 2 years ago

    cracks? it doesn't even exist. we figured this out a long time ago.

  • source
  • [–] 27 points 2 years ago* (last edited 2 years ago) (2 children)

    I feel like a draft landed on Tim's desk a few weeks ago, explains why they suddenly pulled back on OpenAI funding.

    People on the removed superfund birdsite are already saying Apple is missing out on the next revolution.

  • source
  • hideshow 2 child comments
  • [–] 20 points 2 years ago (3 children)

    I hope this gets circulated enough to reduce the ridiculous amount of investment and energy waste that the ramping-up of "AI" services has brought. All the companies have just gone way too far off the deep end with this shit that most people don't even want.

  • source
  • hideshow 3 child comments
  • [–] 17 points 2 years ago*

    Here's the cycle we've gone through multiple times and are currently in:

    AI winter (low research funding) -> incremental scientific advancement -> breakthrough for new capabilities from multiple incremental advancements to the scientific models over time building on each other (expert systems, LLMs, neutral networks, etc) -> engineering creates new tech products/frameworks/services based on new science -> hype for new tech creates sales and economic activity, research funding, subsidies etc -> (for LLMs we're here) people become familiar with new tech capabilities and limitations through use -> hype spending bubble bursts when overspend doesn't keep up with infinite money line goes up or new research breakthroughs -> AI winter -> etc...

  • source
  • [–] 15 points 2 years ago

    They predict, not reason....

  • source
  • [+] 9 points 2 years ago* (last edited 1 year ago)
    load more comments
    view more: next ›