My colleagues in math are now frightened about the (very expensive) mathematical theorem proving ability of these AIs, and many of them really do think that if they can do math, they can do all cognitive tasks. Running a store like this should be so easy! Every single conversation about AI with them has become more frustrating. They are so confused when I still say that the AI companies will die a painful death. When I give my usual points about their expense and their failures in other domains, I am given the usual spiel of "it'll get better in other areas" and "it'll get cheaper".
Unlike them, I have actually been paying attention to this stuff from the beginning. What they think is going on is AI solving math first and shortly getting around to all the other stuff, but what I've seen is that AI labs had already tried all the other stuff first and only managed to win the booby prize of theorem proving, which doesn't pay the bills. And what's the point of spending thousands or millions to output random blobs of Lean that technically compile if there is no one around to bother making sense of them?
One example I gave is when Anthropic vibe coded an entire C compiler from scratch back in February, which turned out to be a pile of shit. I've said that if AI had made similarly rapid progress on software engineering, we would have seen Anthropic continue to put out these demonstrations, and they would have become truly high quality. They would release a compiler more efficient than gcc one week, and a browser better than Chrome the next. (OpenAI's actual attempt at a browser didn't go so well.) And if they could do this, they would actually have a shot of making money!
If they could do this, they would have already. The theorem proving stuff actually works (for certain things, in certain ways, at enormous expense), and look at how OpenAI and Anthropic do not hesitate to snipe mathematicians for results rather than being content as tool vendors. But lately I haven't heard of any software demonstrations. Silence is much louder than noise. More Millennium prize problems bashed with tens of millions in compute costs are not going to change my mind very much.
The counterargument I got was that AI can already one-shot most programming tasks and I shouldn't be cherry-picking the failures. I am far too tired to argue at this point.