top 50 comments

sorted by: hot top controversial new old
[–] 106 points 1 day ago (5 children)

the price of thought

Oh fuck off.

  • source
  • hideshow 5 child comments
  • [–] 38 points 1 day ago (3 children)

    These fuckers would sell us air if they could.

  • source
  • parent
  • hideshow 3 child comments
  • load more comments (1 reply)
    [–] 14 points 1 day ago (1 child)

    The other day I saw a video talking about this new innovation on LLM inference side of things where they keep some more used weights in RAM and others less used on disk. I always suspected from the sample code I stumbled upon on the IA world that should be extreme opportunities for optimizations. But I cannot stress it enough how dumb the LLM world is where the basics of implementing an LRU cache is pass of as some big innovation. Like any half competent comp-sci or comp-eng professional know about the basics of mitigating this basics bottlenecks like "the data does not fit on available RAM", "The disk is slow", etc.

    So is not surprising that now that it seems that the "powerfulness" of this LLMs is starting to plateau that we would start to see some improvement in performance/resource utilization and hence running costs.

  • source
  • hideshow 1 child comment
  • [–] 7 points 1 day ago*

    That works for "Mixture of Experts" models. These are basically models with distinct sets of weights and only a subset of them will be used on any particular query. The rest can sit on a disk.

    It doesn't work for dense models, where every weight is used all the time. There's nothing inactive so a cache has nothing to exploit.

  • source
  • parent
  • [–] 22 points 1 day ago (3 children)

    I wonder how they plan to match "prices are in free fall" to "the AI industry will have to make trillions a year in order not to go bust".

    On the other hand, "prices in free fall" might be the answer they got from AI...

  • source
  • hideshow 3 child comments
  • [–] 18 points 1 day ago (2 children)

    At this point the cloud model firms are basically banking on making a superinteligence before anyone else and taking over the planet, otherwise they go bankrupt. I wish I was kidding.

    Nvidia wins either way though, local models, cloud models, shovels always sell. So they have that overvalued but still realistic bedrock to build houses of cards on.

  • source
  • parent
  • hideshow 2 child comments
  • [–] 6 points 23 hours ago (1 child)

    Nvidia wins unless some other company starts selling cheaper, faster, more efficient matrix multiplication machines.

    I've read some articles about radically different inference architectures that may tilt the scales, but I know this is wishful thinking because I really would like Nvidia to fail badly.

    Linus_nvidia.gif

  • source
  • parent
  • hideshow 1 child comment
  • [–] 1 point 5 hours ago*

    Plenty have tried, all have fallen over flat on their face when it comes to actually providing usable drivers or production at scale. AMD's still completely half assing it even today and Intel's OneAPI and Vino is a bloated joke.

    But yes I would love a future where AMD finally hires an actual software team.

  • source
  • parent
  • [–] 40 points 1 day ago* (last edited 1 day ago) (6 children)

    “Look how low our cost of inference is”

    “Pay no attention to the marketing budget that exceeds Coca-Cola’s for a small fraction of their revenues”

    What technical, fundamental reason is there for the crash in price? The article just accepts the MSRP as fact. It’s established fact that retail prices can be dropped below cost in order to establish market dominance. The cost of training can indeed be spread over time but it’s not spread across enough time (between model releases)The inference cost doesn’t actually drop in reality.

  • source
  • hideshow 6 child comments
  • load more comments (5 replies)
    [–] 66 points 1 day ago (3 children)

    How about DRAM? Show me some AI price crashes leading to DRAM price crashes. That's what we all want, I think.

  • source
  • hideshow 3 child comments
  • [–] 54 points 1 day ago* (last edited 1 day ago) (2 children)

    This still applies.

    Surveillance fascism needs the RAM, compute, and storage to implement totalitarianism and autonomous killing machines that won't refuse to genocide the proles once they realize climate change is significantly worse than advertised, and our "democracies" are an illusion controlled by a big club of plutocrats.

    When the AI bubble pops, the US government will bail out all the tech companies, and all of the debt will be transferred to the working class via our retirement account losses and inflation; no different to the trillions in "forgiven" corporate loans central banks around the world printed during covid. The working class will essentially pay for the nazi big brother and nazi skynet that enslaves them.

    Thanyou for coming to my conspiracy theory ted talk.

  • source
  • parent
  • hideshow 2 child comments
  • [–] 10 points 1 day ago (1 child)

    I’ve been starting to think that once the bubble begins to pop, the US will bail out the banks by buying up their loans to hyperscalers, and they’ll start writing giant surveillance contracts to AI companies to compensate for the lack of demand. The rich get richer and the resulting stock corrections will end up hurting regular people most

  • source
  • parent
  • hideshow 1 child comment
  • [–] 3 points 1 day ago

    Yeah, whether they buy the excess hardware in a fire-sale, nationalize open AI/Anthropic into the DoD on nat sec grounds, or just sign several hundred billion dollar contracts with them, we're just splitting hairs. The threat and their intentions remain the same.

  • source
  • parent
  • [–] 41 points 1 day ago (2 children)

    Extremely interesting… so the depreciation of older models is extreme, while new models are constantly presented. And all the while no AI company is making any profit. This whole story is bonkers

  • source
  • hideshow 2 child comments
  • [–] 17 points 1 day ago*

    The catch is

    the cost of a given level of AI performance

    And more interestingly the article itself says this

    And to this extent, when comparing price drops for AI to drops for other technologies for which we have price series, we are comparing apples and oranges.

    Even then, they decided to make it the headline. This is just like LLM bros doing things they don't know anything about. Absolute garbage.

  • source
  • parent
  • [–] 26 points 1 day ago (2 children)

    So in essence the price that the market will bare for the cost of AI usage is significantly lower than what the big AI companies would like it to be (in order to pay back their ever growing debts), which means there is a possibility they might never achieve profitability on their own (without some external/governmental strong-arming)

  • source
  • hideshow 2 child comments
  • [–] 19 points 1 day ago (3 children)

    That's cool and all, but when can I buy RAM again?

  • source
  • hideshow 3 child comments
  • [–] 13 points 1 day ago (2 children)

    In 4 years or never. The latter probably being the most likely, since they are not just keeping RAM from you for AI purposes. They don't want you to own you own hardware anymore, so they just simply stop manufacturing consumergrade hardware.

  • source
  • parent
  • hideshow 2 child comments
  • load more comments (1 reply)
  • [–] 28 points 1 day ago

    It's funny watching them rig the system and simultaneously keep shooting themselves in the dick.

    Nvidia - desperate not to lose business as they're now ~93% dependent on AI sales, so they keep 'investing' in OpenAI, Anthropic, etc.. Who turn around and of course immediately buy Nvidia AI chipsets.

    OpenAI and Anthropic - panicking that investors will realize their IP is worth nothing (what we've said all along) as they are overtaken by much cheaper models, so they lower their pricing drastically - can't risk losing that market share*.

    *market share is irrelevant really, there is no first-to-market winner in AI, but you cant lie to idiots investors for another 16 rounds of funding to 2030 unless you can show userbase growth to them.

    Really hard to keep propping up the con when barely anyone is paying.

    Fingers crossed for horrible things to happen to then soon.

  • source
  • [–] 2 points 23 hours ago* (last edited 22 hours ago)

    The cost decline for a given level of performance does tend to slow over time

    Even in their tests, there are big drops in 2026 models, and recent ones.

    A much more comprehensive and easy test is to follow this benchmark suite (AAi). Its a mix of medium/hard benchmarks, with bias for agentic coding. (sorry for url dump) https://artificialanalysis.ai/?endpoints=openai_gpt-5-2-codex%2Cazure_kimi-k2-thinking%2Camazon-bedrock_qwen3-coder-480b-a35b-instruct%2Camazon-bedrock_qwen3-coder-30b-a3b-instruct%2Ctogetherai_minimax-m2-5_fp4%2Ctogetherai_glm-5_fp4%2Ctogetherai_qwen3-next-80b-a3b-reasoning%2Cgoogle_gemini-3-pro_ai-studio%2Cgoogle_glm-4-7%2Cmoonshot-ai_kimi-k2-thinking_turbo%2Cnovita_glm-5_fp8&models=mimo-v2-5-pro%2Cgpt-6-sol-low%2Cgemini-3-5-flash-minimal%2Cgpt-5-6-luna-low%2Cclaude-fable-5-1-low%2Ckimi-k2-6%2Ckimi-k2-6-non-reasoning%2Cmimo-v2-0206%2Cglm-5-3-flash%2Cgpt-6-luna-xhigh%2Cglm-4.5%2Cgpt-5-5%2Cclaude-opus-5-low%2Cmimo-v2-5-0424%2Cclaude-sonnet-5%2Cmimo-v2-6-pro%2Cclaude-opus-4-5-thinking%2Cminimax-m3%2Cgpt-6-astra%2Cclaude-opus-5-5%2Cqwen3-8-27b-non-reasoning%2Cgpt-6-luna-medium%2Cclaude-fable-5-1%2Cgpt-5-6-luna%2Cmimo-v2-5-pro-non-reasoning%2Cminimax-m2-7%2Ck2-horizon-mova-36b-a4b%2Cclaude-opus-4-6-adaptive%2Cgpt-5-6-luna-medium%2Cgemini-3-1-flash-lite-preview%2Cgpt-5-4-pro%2Cgrok-4-3-medium%2Cnvidia-nemotron-3-super-120b-a12b%2Cgpt-6-luna-non-reasoning%2Cqwen3-8-27b-medium%2Cgpt-5-5-medium%2Cdeepseek-v4-pro-0424-non-reasoning%2Cgpt-6-luna%2Cgemini-3-8-flash-medium%2Cnvidia-nemotron-3-nano-30b-a3b-reasoning%2Cgrok-4-5%2Cgemini-3-flash-reasoning%2Cgpt-6-luna-high%2Cqwen3-8-27b-low%2Cmimo-v2-flash%2Cdeepseek-v4-pro%2Cqwen3-6-35b-a3b%2Cgemini-3-8-flash%2Cclaude-4-5-sonnet-thinking%2Cllama-4-maverick%2Cgrok-4-3%2Cclaude-opus-4-8%2Cmuse-spark-1-3%2Cqwen3-8-flash-next%2Cgpt-6-astra-low%2Ckimi-k2-5%2Cgpt-5-4%2Cqwen3-8-27b%2Cqwen3-8-max%2Cgpt-5-5-high%2Cclaude-sonnet-5-5%2Chy3%2Cclaude-opus-5%2Cgpt-5-4-mini%2Ck2-horizon-375b-a23b%2Cgemini-3-1-pro-preview%2Cgpt-6-luna-low%2Cgpt-6-sol%2Cgrok-4-6%2Cclaude-4-1-opus-thinking%2Cglm-5-3%2Cgpt-6-sol-non-reasoning%2Cclaude-sonnet-5-non-reasoning%2Cgpt-5-6-sol%2Cgpt-5-6-sol-xhigh%2Cdeepseek-v4-1-flash%2Cgemini-3-7-flash-low%2Cclaude-opus-5-5-xhigh%2Cclaude-sonnet-4-6-adaptive%2Cglm-5-2-non-reasoning%2Cclaude-opus-5-5-medium%2Cgpt-oss-120b%2Cclaude-opus-5-5-high%2Cglm-5-2%2Ckimi-k3%2Cgrok-4-3-low%2Cdeepseek-v4-flash%2Cclaude-opus-5-medium%2Cgemini-4-argon&agents=claude-code-fable-5-1-max-with-fallback%2Ckimi-code-cli-kimi-k3%2Cgrok-build-grok-4-7-xhigh%2Cmuse-code-muse-spark-1-3-max%2Cantigravity-sdk-gemini-3-8-flash-high%2Cdevin-fusion-cli-gpt-6-astra-xhigh-swe-2-medium%2Cclaude-code-qwen3-8-max%2Cdevin-fusion-cli-claude-fable-5-1-xhigh-swe-2-medium%2Ccodex-gpt-5-6-sol-max-reasoning-effort-max%2Cclaude-code-opus-5-max%2Ccodex-gpt-6-astra-max-reasoning-effort-max%2Copencode-glm-5-3-reasoning-effort-max%2Cgrok-build-grok-4-6-xhigh%2Ccodex-deepseek-v4-pro-0813-max%2Ccodex-deepseek-v4-flash-0731-max%2Cmuse-code-muse-spark-1-3-xhigh&coding-agents=cost&model-creators=zai%2Cgoogle%2Calibaba%2Copenai%2Canthropic%2Cxai%2Cmeta%2Cmistral%2Cdeepseek%2Cstepfun%2Cthinking-machines%2Ckimi%2Cifm%2Cminimax%2Cxiaomi&releases=claude-opus-5-5%2Cclaude-fable-5-1%2Cgpt-6-astra%2Cgpt-5-6-luna%2Cmuse-spark-1-3%2Cgrok-4-7%2Cmimo-v2-6-pro%2Cglm-5-3%2Cgemini-3-8-flash%2Cdeepseek-v4-1-flash%2Cminimax-m3&capability-models=gemini-3-5-flash-lite%2Cstep-5%2Cinkling%2Cmimo-v2-6-flash%2Cglm-5-3-flash%2Cmimo-v2-6-pro%2Cminimax-m3%2Cgpt-6-astra%2Cclaude-opus-5-5%2Cclaude-fable-5-1%2Cgpt-5-6-luna%2Cmuse-glimmer%2Cgrok-4-7%2Cnvidia-nemotron-3-ultra-550b-a55b%2Cgpt-6-luna%2Cgemini-3-8-flash%2Cmuse-spark-1-3%2Cqwen3-8-27b%2Cqwen3-8-max%2Cclaude-sonnet-5-5%2Cclaude-opus-5%2Ck2-horizon-375b-a23b%2Cgpt-6-sol%2Cglm-5-3%2Cgpt-5-6-sol%2Cdeepseek-v4-1-flash%2Cmistral-medium-3-5%2Ckimi-k3%2Cqwen3-8-flash-next&cost=evaluation-breakdown

    Opus 4.8 (may 2026), by far best model at the time, gets equaled by deepseek 4 pro (july 31) and 4.1 flash (sept 10th) at less than 1/100th the cost. MiMo 2.6 pro (sept 21) is even cheaper and beats opus 4.8 scores by a wide margin. sonnet 4.6 max to gpt luna high is also a 99% cost drop in a short time for lower performance level models.

  • source
  • [–] 4 points 1 day ago (1 child)

    Moore's Law hasn't been applicable for years.

  • source
  • hideshow 1 child comment
  • [–] 11 points 1 day ago

    Kinda sus that the cost of electricity stops at 1973.

  • source
  • [–] 17 points 1 day ago (6 children)

    Because those other technologies are infinitely more useful.... So obviously governments and private equity you're going to invest in AI. Makes perfect sense to me.

    God I'm so tired

  • source
  • hideshow 6 child comments
  • load more comments (6 replies)
    [–] 6 points 1 day ago

    So revenue is falling. The only way the bubble grows is forcing this shit into even more places?

  • source
  • [–] -2 points 18 hours ago

    Oh yeah "intelligence" is cheap AF...

  • source
  • [–] 2 points 1 day ago

    That's cause you burnt money equivalent of gdp of Spain in that time frame, and I am being generous here.

  • source
  • [–] 7 points 1 day ago (1 child)

    So also less revenue for the AI companies?

  • source
  • hideshow 1 child comment
  • load more comments (1 reply)
    [–] 2 points 1 day ago* (4 children)

    All that R&D, and they need to squeeze efficiency out more at this stage. Faster and more specialized is better than big chingus do everything, and means we can hopefully stop spending all that money on making more datacenters.

    Personally, I don't use the big chingus models at all... I much prefer local gen if I use it for things. Much better privacy that way, and I'm not throwing money at these jokers.

  • source
  • hideshow 4 child comments
  • [–] 1 point 1 day ago (3 children)

    Are there some local models that one can use that are not giving our data away and on par with chat GBT?

  • source
  • parent
  • hideshow 3 child comments
  • [–] 2 points 21 hours ago* (2 children)

    Ollama, Open WebUI, and Qwen, or one of the edits of it. The latest Qwen 26B is in a very good delta of smart enough to know how to help with stuff (I use it as a second opinion/initial proofread code checker) and small enough to run on a macbook or a gaming pc, and Open WebUI is the self host ChatGPT-like interface that works with Ollama, the turn key ai model runner backend. Couple that with your own searxng instance self hosted to give it the ability to search the web for answers if you want to give it that. Run it all in a virtual machine, of course.

    For media content, wan2gp (wan2 for the gpu poor) has an all in one stop shop for self hosted audio, image, and video. Just beware people don't like generated stuff in prod, but its useful for quick concepts, or to do certain types of edits of existing stuff.

    There's also tts and speech recognition via the whisper models, which sounds a lot better than the speak and spell voices from the past (unless you enjoy the degeneracy of spamming JOHN MADDEN JOHN MADDEN in the Steven Hawking voice) and can transcribe words a lot faster (realtime) and more accurately than in the past. The best ones run via Vulkan and thus work on any card that supports Vulkan.

    If you run Nextcloud, you can also connect Ollama to it to get the Nextcloud AI working.

  • source
  • parent
  • hideshow 2 child comments
  • [–] 1 point 1 hour ago (1 child)

    Great input, thank you I appreciate it. Next off for me is learning how to do a virtual machine! I do have ollama currently on a computer connected to a speech to text app. And I find that very useful to have

  • source
  • parent
  • hideshow 1 child comment
  • [–] 1 point 1 minute ago*

    Virt-manager for a gui, qemu-kvm and libvirtd for the hypervisor. You will need to pass thru a graphics card, which makes it unavailable for the host machine entirely. I use debian stable for the vm, no gui, just ssh to manage it.

  • source
  • parent
  • [–] 6 points 1 day ago

    So instead of 1$ return for 3$ spent it's now 4$ or 5$ spent.

  • source
  • [+] 0 points 20 hours ago (1 child)
  • load more comments
    view more: next ›