you are viewing a single comment's thread
view the rest of the comments
[–] 9 points 4 days ago (3 children)

Someone pointed out that on some tests, Astra is actually a regression from previous models. Totally AGI guys

  • source
  • hideshow 3 child comments
  • [–] 6 points 4 days ago

    Yeah sloppers I know went from “see it proved new math! that totally clears us of plagiarism accusations we get when we repost that someone ai generated a frogger and a WoW clone” to “tried to use it for coding and it didn’t do a good job”.

    The math stuff is mostly using Lean proof verifier and all that, by the way, the only use of LLM that they found that is actually legit because it doesn’t matter if its slopping, and you want your attempts randomized.

  • source
  • parent
  • [–] 3 points 4 days ago

    Benchmaxxing for one set of benchmarks can actually degrade performance on other benchmarks. They've plateaued for a while, all they can do is scale inference compute up and down (at logarithmically poor rates of exchange) and trade performance in one area for performance in others

  • source
  • parent