well messages are clearly not stateless (otherwise there would be no context), but in general yes the issue is not the lack of capability, it's the complete unawareness of it and the insistence on lying about it.
THIS time it is ridiculously obvious but what if it does this after checking a very large data set where there would be no (good) way to verify its answer?
This is why Ai, in it's current form, is basically useless. If you cannot trust it NOT to lie, and must/should verify everything yourself, you might as well skip the useless step of asking