It can legitimately find some hard to find bugs but it can quite often come up with stuff for which fix would complicate the code substantially without and practical gains (bugs under very hard to encounter situations etc). So it feels like the biggest problem will arose for setups where people drive this fully autonomously where bug fixes also are implemented autonomously. What would be interesting is whether if this improves when they task a seperate agent to assess practical relevance vs increased code complexity affecting code maintainability and increased interaction between subparts. Both maintainability and practicality however are "long horizon" tasks hard to evaluate without real long term experience.
you are viewing a single comment's thread
view the rest of the comments
view the rest of the comments
replies: