There are constantly new techniques being developed to speed up inference or reduce model size post training. There's about a hundred different levers to pull that play on inference cost, and some of them don't have much of an impact on quality.
In the end, it's probably simply because of competition. I don't understand why everybody assumes they are running these at cost API wise.
Edit: here's a chart from ars technica. The cost of revenue is clearly lower than the actual revenue. They aren't running inference at a lost or it would be higher. They aren't profitable because they are spending all that money on capturing the market as quick as they can. For fucks sake, that includes all the free accounts as well running inference. How can anyone think a paying customer is getting it at less than cost?
And yes, there have been advancement made. The cost of inference isn't some static number that can never go down. Stop believing them when they tell you there's no profit to be made. It's the same playbook as a dozen other companies, they literally get rewarded for it when it comes tax time.
