You might want to check up on the current controversy surrounding OpenAI's math "achievements".
I'm aware of the controversies, this benchmark isn't about making new proofs on previously unsolved problems, it's whether it can answer complex math problems, which it's getting better at.
By my reckoning, the difference between Mythos and Opus is smaller than the difference between Opus and Sonnet. Same with the difference between GPT 5.5 to 5.6 is smaller than the difference between GPT 4 to GPT 5.
Do you have any benchmarks or data to back this "reckoning"