What this does
The photographs on a full disk are rarely there because anyone chose them. They accumulated — a phone that backs up automatically, a camera emptied twice a year, a folder called Import that nobody has opened since 2019. The size of the problem is measured in gigabytes and the individual files are not interesting.
So the shape of the work is a pass over a directory tree, not a session with an image. One rule, applied in order, with the folder structure carried through to the output so that what comes back is recognisably the same library and not a flat bag of four thousand files called IMG_ something.
Setting the rule
- Standard quality is the starting position for an archive. Strong is a reasonable choice for a folder you are keeping out of completeness rather than affection, and a poor one for the photographs you would miss.
- Capture information stays by default here. The date a photograph was taken is frequently the only thing that makes a 2013 folder navigable at all, and it costs a few kilobytes a file.
- The long edge cap is where the gigabytes actually are, and it is irreversible from the output. If you are going to use it on a library, use it on a copy first and look at the result on the largest screen you own.
- Leave the skip rules on. They are why the total at the end is a number and not an estimate.
What it will open
A photo library is never only photographs. Video from the same phone, RAW files beside their JPEG previews, sidecar files from whatever catalogued them once, and the occasional PDF someone dropped in there. All of it is counted, none of it is touched, and the count is shown before the queue starts so the shape of the folder is not a surprise at file three thousand.
Photographs from an iPhone are commonly HEIC, and this build has no HEIC decoder. Rather than counting them as photographs and then skipping every one of them, the pickers do not offer the format, the count leaves them out, and the summary above the queue says how many are in the folder and that they are being left alone — before anything starts.
Questions about long jobs
- Does this replace my library?
- No, and it cannot. The results are written into a separate archive with the folder structure copied. Deleting anything is a decision you make afterwards, with the manifest in front of you, on your own file manager.
- How do I check a job before I delete anything?
- The manifest is a CSV with one line per file: its path, what it weighed, what came out, the outcome and the reason. Open it in a spreadsheet, sort by outcome, and read the skipped and failed rows. If those rows account for everything missing from the archive, the job is complete.
- How long does a few thousand photographs take?
- It is a function of megapixels rather than files, and of one decode at a time by default. A folder of phone photographs moves quickly; a folder of 45-megapixel raw exports converted to JPEG does not. Leaving it running is the expected way to use this, which is why the queue survives the tab closing.
- Is a second pass over an already-compressed library worth anything?
- Rarely, and the skip rule will tell you so. Files that came from a previous run re-encode larger, get thrown away and are counted as skips, so a pointless second pass reports itself as a pointless second pass instead of quietly costing quality.