I'd be curious to see how they defined "harmful" in their test, since it looks like Claude code's technically-not-yolo mode is looking for "actions that are irreversible, destructive, or targeted outside of your environment" which covers a lot of normal actions, especially depending on how you define which resources are in "your environment".
And yes, I'm sure that the fact that the flow of operations they're outline will massively increase the number of times that the system automatically retries things, this increasing token usage, is going to be fully accounted for when they review the financials on how many tokens people are paying for.