
🤔

🤔
But it was incorrect at 15,823 tokens per second!


Well now what?
AMD might win the "AI race" with this. Let the other ones fall, then come in with affordable purpose built hardware that's actually useful.
The general service model won't keep working forever, but a chip that's specifically designed to like translate from Spanish to English, or OCR documents, or whatever other specific LLM task you need is a product that can be sold repeatedly.
On the specific chip, I've seen another format called compute-in-memory from Anker.
So it will be interesting to see what comes out of the AI tech race.
I think the GPU model of "a thing you jam data into and get a thing out" is gonna become more common. Baking a network into a chip is more efficient than fpgas too.
I found a YouTube link in your comment. Here are links to the same video on alternative frontends that protect your privacy:
I found a YouTube link in your post. Here are links to the same video on alternative frontends that protect your privacy:
On the road to fully automated luxury gay space communism.
Spreading Linux propaganda since 2020
Rules: