I asked an LLM to generate tests for a 10 line function with two arguments, no if branches, and only one library function call. It's just a for loop and some math. Somehow it invented arguments, and the ones that actually ran didn't even pass. It made like 5 test functions, spat out paragraphs explaining nonsense, and it still didn't work.
This was one of the smaller deepseek models, so perhaps a fancier model would do better.
I'm still messing with it, so maybe I'll find some tasks it's good at.