So I have a problem. I have WAY too many hard drives and backups from decades of computing. I just hoard everything. The issue is now I don't know what I have duplicates of and what I don't. or when I want to format an old external drive for something else, I'm not sure if I actually have a backup of it.

I just want a way to basically compare folders and see what files are missing or what are duplicates. I'm fine with CLI but I may prefer a GUI for this so I'm not deleting files by accident as easily...

I have searched around for this question but all the solutions seem way too complex for me.

you are viewing a single comment's thread
view the rest of the comments
[–] 7 points 18 hours ago

rsync has a --dry-run / -n flag that would tell you which files would be synced. Combined with --archive /-a it runs recursively.

That said, it's assumes file paths are matching between source and destination, so if things are scattered, it might not work well.

If you're really scattered, I'd consider using a simple python script to walk directories and create a json file with results (maybe the sha256 for key and filename, size, absolute path, and created/modified stamps as fields). Then when you hit a duplicate key, you can dump the results and prompt a delete.

Or something. I'm just spit balling here.

  • source