@bruh I don't know, you might be right. I should play with running an LLM on my own GPU, I haven't don that yet. I've been thinking about it though.
As far as Chinese LLM's I'm not sure? I mean, I generally have the philosophy you get what you pay for. Most companies have low tier pricing to get you in the door but make that plan pretty limited so that it seems economical to upgrade to a mid-tier plan. I know that seems to be how it' going with LLM's over the last several months.
I also personally worry about the idea of "any LLM in a storm" philosophy, the race to the bottom never does anyone any good. Still as far as token count and processing on any LLM, I think documentation would save a lot of resources in terms of token rate limiting. A large code base is a lot of tokens for an AI to process.
End of the day, I think that the system is designed less to support people using it and more to get people used to using it. I think the main goal of the AI industry is to replace human workers with AI agents so they can save $$, get rid of HR departments (or at least scale them back significantly), save money on legal representation, taxes, etc...