that ASIC chip prototype is pretty impressive. You can try it on https://chatjimmy.ai/ without an account, ask it to write something big like an essay or guide - literally the longest you'll wait is to get connected to the API, but the answer appears instantly.
Only limitation right now is they put a small llama 8b model on their chip, but it's a prototype and proof of concept of course. I'm sure soon China will print a full Deepseek model on such a chip lol.
Right now there isn't much interest in making AI more efficient to run but yeah there's no reason we won't find advances there. China is already doing a lot to squeeze models into smaller hardware.
I don't run LLMs locally because what I'm limited to is not great (especially context size is limited) but the way things are going we will definitely start to see open options open up, I think. If only because academia requires it.