It's hard to say. My intuition is that LLMs themselves aren't capable of this until some trillion-parameter neural phase transition maybe, and more focus needs to be put on the cognitive architectures surrounding them. Basically, automated hand-holding so they don't forget what they're supposed to be doing, the equivalent of the brain's own feedback loop.
The main issue is executive function is such a weak signal in the data that it would probably have to reach ASI before it starts optimizing for it, so you either need specialized RL or algorithmic task prioritization.