I am not the person you are replying to, but if the accelerometers are sensible enough, the vibration of the voice will be picked up by the accelerometer.
Since the sound we make when talking are periodical, it can probably easier to track that periodicity and reconstruct the sound from there.
It's all my (un)educated guess.