I'm a noob to local LLMs. I want to use an LLM to create Python and Bash functions from well-defined specs written by me. I've used Claude's Sonnet for this mostly.

I'm a light user of LLMs and almost never hit the request limit on Claude.

It seems like cloud LLMs are in a price war right now and last-gen capabilities are bottoming out in cost? Is that correct?

I pay about $0.12 per kWh, probably going up as more AI data centers get built.

Hardware I have:

  • <16GB VRAM - AMD BC-250 APU - "16 GB total with approximately 12 GiB assigned to GPU UMA and 4 GiB left for the OS" (hardware unlocking changes by the day)
  • 8GB RAM - 11th gen Intel laptop with Xe graphics
  • 8GB RAM - Pixel 7a
  • 32 GB DDR3 - 2nd gen intel - doubt this does anything

Hardware I'm eventually selling:

  • 8GB VRAM RTX 3070 + 32 GB DDR5 + 1TB NVMe SSD - AMD Ryzen 5 7600X CPU - putting this here in case it's substantially better than the BC-250

Among this hardware, I should look for a model that fits on the BC-250?

you are viewing a single comment's thread
view the rest of the comments
[–] 4 points 1 week ago

Qwen3.5-35B-A3B is the most powerful model you could run on your sell-eventually desktop, because it is MOE (mixture of experts) you can split its 22GB 4bit variant over both CPU and GPU and achieve workable performance. Use llama.cpp-cuda directly to save overhead, Claude can even help you set up a performance tested start script for your configuration. Hopefully Qwen3.8 will drop soon in a similar variant. On the other hand, deepseek 4 flash on openrouter.ai beats it easily and is energy cost level cheap. I would take that route over smaller models on programming duty on your weaker hardware.

  • source