Gemma 4 12B QAT + faster Qwen 3 TTS fits perfectly fine in 12GB VRAM and can be used to create a fully local voice assistant with ... reconizeable traits ;)
I see Youtubers trying to do similar things every now and then but for some reason their latencies always seem a lot higher than I would expect so if the one I've conjured up (based on Huggingface's Speech-to-Speech project initially) is special somehow then maybe it's of interest to others.
Btw, just to let you know, in case your "computer" model is still to be trained... (as you write in the Readme) There's a bunch of people who make microwakeword models for ESPHome and put them on GitHub: