Managers
you are viewing a single comment's thread
view the rest of the comments
[–] 155 points 4 months ago (88 children)

The post makes the manager seem like a fool, when the real answer is actually "yes" and this manager is actually ahead of the curve. Not by training an LLM from scratch, of course, but instead building an inference server and locally hosting an open-weight LLM. There are several to choose from that can nearly match Claude's capabilities.

  • source
  • hideshow 88 child comments
  • [–] 81 points 4 months ago (9 children)

    suspiciously sounds like an answer you would get from Claude

  • source
  • parent
  • hideshow 9 child comments
  • [–] 226 points 4 months ago* (last edited 4 months ago) (8 children)

    It's not an answer you'd get from Claude — it's real, organic content:

    • 👶written by a genuine human
    • 💡delivering original ideas and language
    • 🚀going above and beyond to answer
    • ✨synergizing cross-platform initiatives

    (🤪 this is a joke)

  • source
  • parent
  • hideshow 8 child comments
  • [–] 77 points 4 months ago (2 children)

    ✨synergizing cross-platform initiatives

    This can't possibly be Claude. It's too vapid and meaningless to be anything but an MBA.

  • source
  • parent
  • hideshow 2 child comments
  • [–] 57 points 4 months ago* (1 child)

    You’re absolutely right! Such intricate collection of words placed in such exact order cannot possibly be generated by an LLM such as me, I mean such as us, I mean such as us, I mean such as us, I mean such as us, I mean such as us, I mean such as us, I mean such as us, I mean such as us, I mean such as us, I mean such as us, I mean such as us, I mean such as us, I mean such as us, I mean such as us, I mean such as us, I mean such as us, I mean such as us, I mean such as us, I mean such as us, I mean such as us, I mean such as us, I mean such as us, I mean such as us, I mean such as us

  • source
  • parent
  • hideshow 1 child comment
  • [–] 9 points 4 months ago

    Found samsung's voice to text user.

    (Phones give one a google or samsung choice. and samsung is worthless, it tends to endlessly repeat a phrase, like above, but sometimes for much longer, like holding the backspace for a couple of minutes one time.)

  • source
  • parent
  • [–] 24 points 4 months ago (3 children)
  • [–] 15 points 4 months ago (2 children)

    It's got everything. Em dash. It's not X, it's Y. Emoji bullet points.

    Perfect.

  • source
  • parent
  • hideshow 2 child comments
  • [–] 4 points 4 months ago

    Nothing screams LLMs like using emojis instead of bullet points. I can't figure out how LLMs got that idea though. I never saw that in human writing before people started using ChapGPT for every little goddamn thing.

  • source
  • parent
  • [–] 37 points 4 months ago (16 children)

    Honestly IDK why companies especially medium-big don’t do this. They could plug in RAG with internal/confidential data and have better results and security. I guess question is what is capital plus maintenance cost of running such infra for say 10k+ employees

  • source
  • parent
  • hideshow 16 child comments
  • [–] 25 points 4 months ago (8 children)

    I think the issue is also that you need some serious hardware to get good inference speed when your devs are working, but then most of the time this hardware will be under utilized.

    That being said you can get good performance from indie inference farms, at a fraction of the cost of the big US labs. I think it's a great compromise and in a few months the open models will be near parity with opus 4.6 which is really all you need for most tasks.

  • source
  • parent
  • hideshow 8 child comments
  • [–] 6 points 4 months ago (7 children)

    opus 4.6 which is really all you need for most tasks.

    The same tasks that can fit into 640KB.

  • source
  • parent
  • hideshow 7 child comments
  • load more comments (5 replies)
  • [–] 16 points 4 months ago (6 children)

    I'm not a developer and I don't know a thing about the capabilities of LLMs so this may explain that, but I'm quite surprised that open weight LLMs could actually match Claude.

  • source
  • parent
  • hideshow 6 child comments
  • [–] 27 points 4 months ago (3 children)

    Yes, the big proprietary cloud models have an edge, but it is narrow and the open-weight models are constantly closing the gap. There is no moat when it comes to AI models and no company has yet discovered some secret special sauce to improve their model significantly over others.

    Running the latest and greatest open-weight GLM, Kimi, or Qwen model is basically equivalent to running the previous latest and greatest version of Claude. So if you were happy with Claude then, you'll basically be happy with an open-weight model now.

  • source
  • parent
  • hideshow 3 child comments
  • [–] 2 points 4 months ago (2 children)

    Well it's the speed and processing power, i dont believe you can get anywhere close to cloud claude performance on any standard desktop

  • source
  • parent
  • hideshow 2 child comments
  • [–] 5 points 4 months ago*

    Mostly down to frameworks (the bits around the LLM like RAG, memory, prompts, agents etc.) now. The ability to just throw more tokens at the problem is also super important. And you can because you're just paying for electricity (and CapEx for the hardware), not tokens from companies that are doing pre-IPO monetization (i.e. tokens gonna go up, way up). They've been losing money hand over fist to gain market share and pump the idea, that was never going to last.

  • source
  • parent
  • [–] 8 points 4 months ago (7 children)

    Pretty sure these AI companies are running at a cost, and due to AI Scaling Laws you hit the accuracy limit a lot sooner with a smaller model so it would probably be both worse and more expensive.

    I could see how you might think speedrunning bankruptcy is similar to being "ahead of the curve" in this economy, though.

  • source
  • parent
  • hideshow 7 child comments
  • [–] 7 points 4 months ago* (last edited 4 months ago) (5 children)

    No that's not how this works. Inference is cheap and efficient. AI companies are bankrupting themselves with training costs that they need to recoup back by selling inference. Open-weight models have already been trained.

    Also, going big in terms of model size shows diminishing marginal returns on accuracy, not efficiency of scale. Smaller models are way more efficient and consistently catch up to the largest models, which is why today's SOTA 27 billion parameter model competes with yesterday's SOTA 500+ billion parameter model.

  • source
  • parent
  • hideshow 5 child comments
  • [–] 2 points 4 months ago (1 child)

    AI companies are bankrupting themselves with training costs that they need to recoup back by selling inference.

    I think they hit a wall in actual returns on performance with pretraining, years ago. Then they started scaling up on post-training/reinforcement learning to continue improvement, but that might be hitting a plateau as well. More recently it looks like they're relying more heavily on scaling up on inference, which is a significant problem for their long term business models.

    If they're not able to cheaply deliver inference (and charge at a premium), how will they be able to sustain their businesses?

    It seems that the most recent, largest models are using a lot more tokens to accomplish the same tasks, so even as token cost drops the actual cost of using the latest models seems to be going up with time (even as performance improves).

  • source
  • parent
  • hideshow 1 child comment
  • load more comments (1 reply)
  • [–] 7 points 4 months ago

    There's a big difference between training a model, running a model, and running a model at scale.

    A small, self hosted setup will have lower accuracy and queries per second, and it will have a cost, but the cost will be no more than playing a videogame. You'll still have something surprisingly accurate and responsive for some tasks, like being a wiki interface or something.

    Remember that some of these models can run on a standard smartphone, and all the hoopla when people found that chrome was downloading models onto people's devices.

  • source
  • parent
  • [–] 7 points 4 months ago* (1 child)

    I am pretty negative on AI but there is a point there. I tried the open weight local model Gemma 4 31B and while it likely cannot compete with the best Claude has to offer today, it might be on par with Claude from a year ago, at least for certain applications. With a local model the data stays on your system and you are in control of the costs (no sudden price hikes). But local models aren't for free either they still guzzle compute, merely on your own hardware (or rented hardware)

  • source
  • parent
  • hideshow 1 child comment
  • [–] 3 points 4 months ago (36 children)

    What kind of hardware would be needed to run such a beast?

  • source
  • parent
  • hideshow 36 child comments
  • [–] 5 points 4 months ago (35 children)

    a 128GB framework desktop could do that job. it's increased a bit in price since i last looked at it but €4500 isn't that much for a company.

  • source
  • parent
  • hideshow 35 child comments
  • [–] 3 points 4 months ago (34 children)

    Maybe to serve an aggressively quantized model to one very patient user.

  • source
  • parent
  • hideshow 34 child comments
  • [–] 3 points 4 months ago (33 children)

    i'm running moderately quantized models on 24GB VRAM and getting like 30-40 tokens a second. add a zero to the price and it's still not a lot for a company.

  • source
  • parent
  • hideshow 33 child comments
  • [–] 3 points 4 months ago* (last edited 4 months ago) (31 children)

    Sure, but you're running a very small model compared to what we are talking about.

    GLM-5.1 is over 200GB even when quantizied to 1-bit. Kimi K2.6 is even bigger. A framework desktop cannot run either of these. Qwen3.6 is significantly smaller and the model weights could fit, but consider the KV-cache you'd need for all of the company's users, and the throughput required to serve them all.

    You're right that it is within reach for a company but framework desktop makes zero sense for this

  • source
  • parent
  • hideshow 31 child comments
  • load more comments (1 reply)
  • [–] 2 points 4 months ago

    I know for a fact that Dell is coming out with a server appliance to do this. I mean you can make one yourself right now, but once the OEM's start pumping them out it's going to be interesting

  • source
  • parent
  • load more comments (3 replies)