Gemma 4 12B QAT + faster Qwen 3 TTS fits perfectly fine in 12GB VRAM and can be used to create a fully local voice assistant with ... reconizeable traits ;)
I see Youtubers trying to do similar things every now and then but for some reason their latencies always seem a lot higher than I would expect so if the one I've conjured up (based on Huggingface's Speech-to-Speech project initially) is special somehow then maybe it's of interest to others.