Single-thread is really hard, we've basically saturated our l1 working set size, adding more doesn't help much. Trying to extend the vector length just makes physical design harder and that reduces clock speed. The predictors are pretty good, and Apple finally kicked everyone up the ass to increase OOO like they should have.
Also, software still kind of sucks. It's better than it was, but we need to improve it, the bloat is just barely being handled by silicon gains.
Flash was the epochal change, maybe we have some new form of hybrid storage but that doesn't seem likely right now, Apple might do it to cut costs while preserving performance, actually yeah I see them trying to have their cake and eat it too.
Otherwise I don't know, we need a better way to deal with GPUs, there's nothing else that can move the needle, except true heterogenous core clusters, but I haven't been able to sell that to anyone so far, they all think it's a great idea, that someone else should do.