▲ 99 ▼ rsync goes AI slop, breaks your backups (pivot-to-ai.com) submitted 3 months ago by dgerard@awful.systems [M] to c/techtakes@awful.systems 28 comments fedilink hide all child comments 36 commits by ‘tridge and claude’ https://www.youtube.com/watch?v=V0EAo9jo-U4&list=UU9rJrMVgcXTfa8xuMnbhAEA - video https://pivottoai.libsyn.com/20260603-rsync-goes-ai-slop-breaks-your-backups - podcast time: 7 min 34 sec
[–] dgerard@awful.systems [S] 11 points 3 months ago (7 children) It literally is slop. It's always correct to call slop slop. permalink fedilink source parent hideshow 7 child comments replies: [–] MoonMelon@lemmy.ml 21 points 3 months ago (6 children) I rewrote the rsync test suite in python from the old shell script design. I did the design for that myself (and I’m really quite pleased with it), but used claude with cross-checks from codex and gemini to do the grunt work. I did not just vibe-code “convert test suite to python”.... I used AI tools to do the grunt work because they are good at that. I reviewed every part of it myself and ran through a huge amount of CI time getting it right If what he claims is true then he's using LLMs for test coverage with significant editing by hand. I hate LLMs, but even I have to admit this seems like one of the few, valid use cases of LLM assisted coding. Unless "slop" has become one of those words that's just lost all meaning. permalink fedilink source parent hideshow 6 child comments replies: [–] dgerard@awful.systems [S] 18 points 3 months ago (3 children) I commend to you jonny's thread on the tests: https://neuromatch.social/@jonny/116666900898570791 It keeps turning out that when you look at the AI output, it's shit. permalink fedilink source parent hideshow 3 child comments replies: [–] MoonMelon@lemmy.ml 16 points 3 months ago (2 children) I don't know anything about rsync aside from as a user, but I am pretty experienced with Python and I admit those tests look really bizarre. If he did "slot machine" code it (a term I wasn't familiar with) then yeah, I agree that's slop. If he didn't, I don't understand why he made these changes. OK yeah, that's a bad sign. permalink fedilink source parent hideshow 2 child comments replies: [–] dgerard@awful.systems [S] 25 points 3 months ago (1 child) every vibe coder insists they're shooting up krokodil responsibly permalink fedilink source parent hideshow 1 child comment replies: [–] arbitraryidentifier@awful.systems 8 points 3 months ago krokodil is such a good analogy goddamn permalink fedilink source parent [–] AnarchistArtificer@slrpnk.net 12 points 3 months ago On one of the BlueSky threads going over over the test code, one of the things they uncovered was some stuff running as root which in no world should be necessary. He may not have just prompted Claude to "convert test suite to python", but there's a lot there that seem like clear red flags in terms of AI slop code. Which is no surprise, really, given that properly proof-reading AI code is often much more labour intensive than just writing the code oneself. It's easy for things like this to slip through the cracks, even if you are trying to check the AI output permalink fedilink source parent [–] diz@awful.systems 11 points 3 months ago* (last edited 3 months ago) It's a perfect example of how "using LLMs for test coverage" can also be harmful. He expected the tests to to prevent introduction of said regressions, probably based on a combination of the quantity of tests and their style (they look like what decent human written tests look like). But the tests are AI slop, and so they give a lot less value per line of code than he expects, hence a significant regression. It is literally useful to call these tests AI slop, and the problem is in part caused by not calling them AI slop, and having consequent inflated expectations. LLMs are not any better at writing tests than at writing other code! It is merely that the bar for tests can, legitimately, be a lot lower (in projects where there would otherwise be no tests at all). Making an exception to calling AI generated tests "slop" is thus counter productive, because it leads people to act as if LLMs are actually better at writing tests than at writing other code, and not just because the bar for tests is frequently very low. edit: actually scratch that I looked at the PR and those tests even look like dogshit and worse than the tests I seen claude write at a workplace that was into vibecoding (which i since quit). permalink fedilink source parent
[–] MoonMelon@lemmy.ml 21 points 3 months ago (6 children) I rewrote the rsync test suite in python from the old shell script design. I did the design for that myself (and I’m really quite pleased with it), but used claude with cross-checks from codex and gemini to do the grunt work. I did not just vibe-code “convert test suite to python”.... I used AI tools to do the grunt work because they are good at that. I reviewed every part of it myself and ran through a huge amount of CI time getting it right If what he claims is true then he's using LLMs for test coverage with significant editing by hand. I hate LLMs, but even I have to admit this seems like one of the few, valid use cases of LLM assisted coding. Unless "slop" has become one of those words that's just lost all meaning. permalink fedilink source parent hideshow 6 child comments replies: [–] dgerard@awful.systems [S] 18 points 3 months ago (3 children) I commend to you jonny's thread on the tests: https://neuromatch.social/@jonny/116666900898570791 It keeps turning out that when you look at the AI output, it's shit. permalink fedilink source parent hideshow 3 child comments replies: [–] MoonMelon@lemmy.ml 16 points 3 months ago (2 children) I don't know anything about rsync aside from as a user, but I am pretty experienced with Python and I admit those tests look really bizarre. If he did "slot machine" code it (a term I wasn't familiar with) then yeah, I agree that's slop. If he didn't, I don't understand why he made these changes. OK yeah, that's a bad sign. permalink fedilink source parent hideshow 2 child comments replies: [–] dgerard@awful.systems [S] 25 points 3 months ago (1 child) every vibe coder insists they're shooting up krokodil responsibly permalink fedilink source parent hideshow 1 child comment replies: [–] arbitraryidentifier@awful.systems 8 points 3 months ago krokodil is such a good analogy goddamn permalink fedilink source parent [–] AnarchistArtificer@slrpnk.net 12 points 3 months ago On one of the BlueSky threads going over over the test code, one of the things they uncovered was some stuff running as root which in no world should be necessary. He may not have just prompted Claude to "convert test suite to python", but there's a lot there that seem like clear red flags in terms of AI slop code. Which is no surprise, really, given that properly proof-reading AI code is often much more labour intensive than just writing the code oneself. It's easy for things like this to slip through the cracks, even if you are trying to check the AI output permalink fedilink source parent [–] diz@awful.systems 11 points 3 months ago* (last edited 3 months ago) It's a perfect example of how "using LLMs for test coverage" can also be harmful. He expected the tests to to prevent introduction of said regressions, probably based on a combination of the quantity of tests and their style (they look like what decent human written tests look like). But the tests are AI slop, and so they give a lot less value per line of code than he expects, hence a significant regression. It is literally useful to call these tests AI slop, and the problem is in part caused by not calling them AI slop, and having consequent inflated expectations. LLMs are not any better at writing tests than at writing other code! It is merely that the bar for tests can, legitimately, be a lot lower (in projects where there would otherwise be no tests at all). Making an exception to calling AI generated tests "slop" is thus counter productive, because it leads people to act as if LLMs are actually better at writing tests than at writing other code, and not just because the bar for tests is frequently very low. edit: actually scratch that I looked at the PR and those tests even look like dogshit and worse than the tests I seen claude write at a workplace that was into vibecoding (which i since quit). permalink fedilink source parent
[–] dgerard@awful.systems [S] 18 points 3 months ago (3 children) I commend to you jonny's thread on the tests: https://neuromatch.social/@jonny/116666900898570791 It keeps turning out that when you look at the AI output, it's shit. permalink fedilink source parent hideshow 3 child comments replies: [–] MoonMelon@lemmy.ml 16 points 3 months ago (2 children) I don't know anything about rsync aside from as a user, but I am pretty experienced with Python and I admit those tests look really bizarre. If he did "slot machine" code it (a term I wasn't familiar with) then yeah, I agree that's slop. If he didn't, I don't understand why he made these changes. OK yeah, that's a bad sign. permalink fedilink source parent hideshow 2 child comments replies: [–] dgerard@awful.systems [S] 25 points 3 months ago (1 child) every vibe coder insists they're shooting up krokodil responsibly permalink fedilink source parent hideshow 1 child comment replies: [–] arbitraryidentifier@awful.systems 8 points 3 months ago krokodil is such a good analogy goddamn permalink fedilink source parent
[–] MoonMelon@lemmy.ml 16 points 3 months ago (2 children) I don't know anything about rsync aside from as a user, but I am pretty experienced with Python and I admit those tests look really bizarre. If he did "slot machine" code it (a term I wasn't familiar with) then yeah, I agree that's slop. If he didn't, I don't understand why he made these changes. OK yeah, that's a bad sign. permalink fedilink source parent hideshow 2 child comments replies: [–] dgerard@awful.systems [S] 25 points 3 months ago (1 child) every vibe coder insists they're shooting up krokodil responsibly permalink fedilink source parent hideshow 1 child comment replies: [–] arbitraryidentifier@awful.systems 8 points 3 months ago krokodil is such a good analogy goddamn permalink fedilink source parent
[–] dgerard@awful.systems [S] 25 points 3 months ago (1 child) every vibe coder insists they're shooting up krokodil responsibly permalink fedilink source parent hideshow 1 child comment replies: [–] arbitraryidentifier@awful.systems 8 points 3 months ago krokodil is such a good analogy goddamn permalink fedilink source parent
[–] arbitraryidentifier@awful.systems 8 points 3 months ago krokodil is such a good analogy goddamn permalink fedilink source parent
[–] AnarchistArtificer@slrpnk.net 12 points 3 months ago On one of the BlueSky threads going over over the test code, one of the things they uncovered was some stuff running as root which in no world should be necessary. He may not have just prompted Claude to "convert test suite to python", but there's a lot there that seem like clear red flags in terms of AI slop code. Which is no surprise, really, given that properly proof-reading AI code is often much more labour intensive than just writing the code oneself. It's easy for things like this to slip through the cracks, even if you are trying to check the AI output permalink fedilink source parent
[–] diz@awful.systems 11 points 3 months ago* (last edited 3 months ago) It's a perfect example of how "using LLMs for test coverage" can also be harmful. He expected the tests to to prevent introduction of said regressions, probably based on a combination of the quantity of tests and their style (they look like what decent human written tests look like). But the tests are AI slop, and so they give a lot less value per line of code than he expects, hence a significant regression. It is literally useful to call these tests AI slop, and the problem is in part caused by not calling them AI slop, and having consequent inflated expectations. LLMs are not any better at writing tests than at writing other code! It is merely that the bar for tests can, legitimately, be a lot lower (in projects where there would otherwise be no tests at all). Making an exception to calling AI generated tests "slop" is thus counter productive, because it leads people to act as if LLMs are actually better at writing tests than at writing other code, and not just because the bar for tests is frequently very low. edit: actually scratch that I looked at the PR and those tests even look like dogshit and worse than the tests I seen claude write at a workplace that was into vibecoding (which i since quit). permalink fedilink source parent