It should be a lot harder than foveated rendering. For foveation you need to know the position of the eye, but for dof you need the shape of the lense inside the eye.
You could try estimating the focus based on what the eyes together point at, but the eyes don't actually focus in perfect match with their stereo vision.
And then if you do manage to track it, you need to undo the eyes lense too, since it's naturally trying to blur the screen away from being in focus.
Currently vr just ignores it and the eye just doesn't focus like it naturally would, but if you wanna have say the ability to focus between two overlapping things with the other blurring out of the way, you need that lense input, so you need the eye to focus, so you need to let the lense squeeze while the screen stays in focus.
Tracking lenses seems so hard I wouldn't be surprised if we get light field displays first. Which sidestep the whole problem and are really the right way to do it. Proper sharpness and blur on a passive screen even without tracking. For a good looking blur you probably only need a handful angular pixels, since you know the exact location the eye is at.