Interesting how he mentions LLMs have a 50% error rate, and predicts how no one will pay for a tool with such a poor rate.

all 3 comments

sorted by: hot top controversial new old
[–] 2 points 1 day ago* (2 children)

Here a comment on this from usernomdeguerre on hacker news (ycombinator.com).

"His" and "GKH" refers to Greg Kroah-Hartmann, who is maintainer of the stable kernel series and the speaker in the video.

I've included a few slides into text that i thought were eye-opening to me:

From his Kernel Recipes 2026 slide on Mythos


  Mythos's 79 vulnerabilities:
  24 - no detail at all "something crashed"
  14 - not a bug at all
  3 - totally made up data
  15 - already fixed in latest release
    - 11 by others
    - 4 by anthropic
  20 - fixes were needed
    - 7 "assume a malicious filesystem image"
    - 2 "assume you can inject a malicious network packet into the middle of the stack"
    - 2 "NOMMU"
    - 6 sctp networking issues for untrusted devices
    - 2 ipv6 minor network issues 
    - 1 gpu driver for local malicious user

GHK called this "10 'real' bugfixes", which to me sounds like there's a wild hype machine around these companies and uncritical parroting of every press release they make that falls apart when you engage the affected real experts.

  • source
  • hideshow 2 child comments
  • [–] 0 points 1 hour ago (1 child)

    It can legitimately find some hard to find bugs but it can quite often come up with stuff for which fix would complicate the code substantially without and practical gains (bugs under very hard to encounter situations etc). So it feels like the biggest problem will arose for setups where people drive this fully autonomously where bug fixes also are implemented autonomously. What would be interesting is whether if this improves when they task a seperate agent to assess practical relevance vs increased code complexity affecting code maintainability and increased interaction between subparts. Both maintainability and practicality however are "long horizon" tasks hard to evaluate without real long term experience.

  • source
  • parent
  • hideshow 1 child comment
  • [–] 1 point 14 minutes ago

    So it feels like the biggest problem will arose for setups where people drive this fully autonomously where bug fixes also are implemented autonomously.

    Completely irrelevant for the kernel and mostly marketing of LLM companies.

  • source
  • parent