So I have a problem. I have WAY too many hard drives and backups from decades of computing. I just hoard everything. The issue is now I don't know what I have duplicates of and what I don't. or when I want to format an old external drive for something else, I'm not sure if I actually have a backup of it.

I just want a way to basically compare folders and see what files are missing or what are duplicates. I'm fine with CLI but I may prefer a GUI for this so I'm not deleting files by accident as easily...

I have searched around for this question but all the solutions seem way too complex for me.

you are viewing a single comment's thread
view the rest of the comments
[–] 1 point 3 hours ago (1 child)

Say you have files in /home/bridgenjoyer/Data, /mnt/Backup, and /mnt/data

So

for d in /home/bridgeenjoyer/Data /mnt/backup /mnt/data
 do
    find $d -type f -print0
  done | xargs -0 md5sum | sort
  • generates a list of all plain files by path name
  • runs md5sum on them, which outputs a checksum as "checksum filename"
  • sorts them by checksum

So, for identical files, you get identical repeated checksum antries.

You can filter out repeated entries (e.g. using awk, a shell script, or python) and extract their name and path.

There are also various tools which deduplicate identical files by hard linking them, if they are on the same file system. These will get the same inode number shown with ls -l. You can find these tools in the Arch wiki.

Then you need a strategy how to organize these identical files. You cold put all files into an archive folder, tidy them up, and only make and keep backups which you don't modify.

  • source
  • hideshow 1 child comment