IMHO it's not the speed. People are patient enough if the result is good. But lets be honest the context windows are damm small to handle local context.
Try to summarize things which are bigger than a email or a very small article.
Try to have a slightly bigger codebase...
And specially this "smaller" local llm's have a much more limited quality by default without additional informations provided.
We also don't wanna talk about the expected prices of DDR5 memory for modern CPU's. So even if you have a AI CPU from AMD or similar most of those PC's won't have 64+GB ram ->
Try of a bigger content window
QWEN3:4b with 256k ctx
