all 26 comments

sorted by: hot top controversial new old
[–] 58 points 1 month ago (15 children)

I think it’s important to note that these models are open-WEIGHT and not open-SOURCE. Open-source would mean we could explore the training data itself. Open-weight models are great, but please let’s not call them open-source.

  • source
  • hideshow 15 child comments
  • [–] 14 points 1 month ago (6 children)

    Yeah, it's like an open-source distro that won't work without an accompanying opaque binary blob.

  • source
  • parent
  • hideshow 6 child comments
  • [–] 4 points 1 month ago (5 children)
  • [–] 2 points 1 month ago (3 children)

    It is more a question of what hardware is like that and if you look up firmware blobs for Linux you'll find a lot, starting with graphic cards.

  • source
  • parent
  • hideshow 3 child comments
  • [–] 2 points 1 month ago (2 children)

    That's the hardware not the distro

  • source
  • parent
  • hideshow 2 child comments
  • [–] 2 points 1 month ago

    Distro was the wrong word. I should have said "package" (though there are also distros that attempt to eliminate blobs within their packages). Examples include manufacturer-issued device drivers, software that allows integration of GPUs, several parts of Android that integrate with the underlying hardware, firmware of various sorts, and I can't be arsed to remember more at the moment.

  • source
  • parent
  • [–] 2 points 1 month ago

    Nemotron Ultra is a open source and already gives DeepSeek R1 performance (admittedly not that good anymore). NVIDIA has open sourced the entire training process, including raw datasets and synthetic data generation. The raw data is mostly curated web crawl data.

  • source
  • parent
  • [–] [S] 1 point 1 month ago (1 child)

    Perfect shouldn’t become the enemy of better. Open Weight AI isn’t Open Source AI, but it’s a huge improvement over Closed AI where you get neither the weights nor the training pipeline.

  • source
  • parent
  • hideshow 1 child comment
  • [–] 3 points 1 month ago

    But to put it another way, what would you say to stalling tactics on the climate? What we need is “perfect” and they are only offering us bottom line-friendly “better”?

    “X-1” doesn’t solve the equation “x-1(x/1000),” know what I mean?

  • source
  • parent
  • [–] 1 point 1 month ago (4 children)

    The EU's AI Act calls them open source, so that ship has sailed.

  • source
  • parent
  • hideshow 4 child comments
  • [–] 27 points 1 month ago* (last edited 1 month ago)

    There are many dialects of self-serving corporate bullshit. It's interesting to occasionally run across a new one.

  • source
  • [–] [S] 5 points 1 month ago (7 children)

    That’s either incredibly idealistic or a very smart strategy: Open models grow ecosystems faster than walled gardens, but compute still decides who gets to play at the highest level. If AGI really becomes shared infrastructure, the moat shifts from models to chips, data and execution.

  • source
  • hideshow 7 child comments
  • [–] 10 points 1 month ago (4 children)

    A bit of both though. In house models don't need the power and process of a small city, they need the power to power what their business needs. Also it would seem to me that has a huge advantage on the whole as a model would be more efficient, if it's focused on the companies needs rather than being a jack of all trades, from poetry to law to code. Cut out the being everything to everyone and you can do far more with far less hardware.

  • source
  • parent
  • hideshow 4 child comments
  • [–] [S] 5 points 1 month ago (2 children)

    This. The frontier race is about building the smartest generalist, but most companies don’t need that. They need a specialist tuned to their workflows. Narrower scope means smaller models, lower costs, faster inference and often better results. General models become the foundation, not the finished product. It’s still all to play for! 💪😎✌️

  • source
  • parent
  • hideshow 2 child comments
  • [–] 2 points 1 month ago (1 child)

    Agreed, However I also wonder how or if LLMs could help with what we don't know we need. With the processing pace of LLMs and the massive context the can processes at once. It seems like it could make the broader connections that are usually the harder to find more accessible.

  • source
  • parent
  • hideshow 1 child comment
  • [–] [S] 2 points 1 month ago

    I think that’s probably the most underrated use case. We keep treating LLMs like faster search engines, but their real value may be surfacing relationships humans never think to test.

  • source
  • parent
  • [–] 3 points 1 month ago

    Exactly. At work, we're already looking at getting our own compute for an open model for some automatons we built. Google's models are alright, but they are retiring them too fast. So we want to get off that treadmill. Plus a local model gives us opportunities to experiment with LoRAs and techniques like what cactus hybrid did with a small head predicting certainty (https://news.ycombinator.com/item?id=49010782). Plus, the big players are removing sampling options thinking they know what's best when it is obvious they don't.

  • source
  • parent
  • [–] 3 points 1 month ago (1 child)

    Remember when OpenAI seemed idealistic? To most people, anyway. This is the same. DeepSeek has already been caught lying to inflate hype. They‘re grifters like all the other AI companies.

  • source
  • parent
  • hideshow 1 child comment
  • [–] [S] 4 points 1 month ago

    OpenAI is a good reminder that mission statements and business incentives don’t always stay aligned. That said, I’d judge DeepSeek less by the rhetoric and more by what they actually release. Hype is cheap: open weights, reproducible results and useful tools are much harder to fake.

  • source
  • parent
  • [+] [S] -1 points 1 month ago