Archiving & preservation workflows
The demand behind "archiving and preservation" is never tidy. It surfaces when a WD MyBook finally dies, when a Google Photos and iCloud export leaves someone with hundreds of thousands of images spread across old cameras, external drives, and a NAS, when a 34 GB Blu-ray rip needs splitting into episodes, or when a DataHoarder starts burning cold data to BD-R with 7-zip and parity because more hard drives cost too much. The files are large, private, often irreplaceable, and usually already a mess: loose ZIPs, missing metadata, half-supported formats, and duplicates on top of duplicates.
The people asking these questions on r/DataHoarder, r/pdf, r/software, and r/WindowsHelp are not shopping for another dashboard. They want to run one concrete operation — dedupe, convert, compress, merge, verify — where the file already lives, check the result, and never hand a private archive to whichever converter ranked first in search. This page collects the archiving and preservation questions those threads keep raising and answers each one directly, whether or not you ever install a desktop tool. The through-line is control: keep the bytes local, preserve the originals, and make every step repeatable so the same cleanup doesn't have to be rebuilt from scratch next month.
How do I organize hundreds of thousands of photos and videos across drives, cloud exports, and a NAS without the tool crashing during import?
This is the single most common version of the archiving problem. One person has hundreds of thousands of photos pulled from Google Photos, iCloud, old cameras, external drives, and a NAS, and every free organizer they try crashes partway through the import. Another has years of images, PDFs, and videos on old removable media and just wants a no-cloud path to consolidate them. A third is staring at disorganized drives, some of them failing, plus huge transfers still in flight — and an unknown number of duplicates and corrupted files in the pile.
The mistake that causes the crashes is loading the entire library into one catalog app that builds a full preview index in memory before it does anything useful. At half a million files that's a memory wall, not a workflow. The reliable approach is to work in place, folder by folder, and lean on a few load-bearing operations: duplicate detection by content hash rather than filename (so renamed copies still match), finding the largest files and empty folders to reclaim space and see what you actually have, and consistent batch renaming so the collection has a predictable structure before you back it up. Keep the originals, act on a copy or with a dry-run preview, and verify a sample of outputs before deleting anything — with irreplaceable media, the verification step is the whole job, not a nicety.
How 1FileTool handles this: it runs locally against a folder you point it at, so nothing has to import into a catalog or leave the machine. Use Find Duplicates to dedupe by content, Analyze Disk Usage and Find Large Files to see where the space went, and Batch Rename to impose a consistent naming scheme across the whole library. See the full File Tools set.
I've got a folder full of loose files, random ZIPs, and old archives with no metadata — how do I make sense of it?
A concrete example from these threads: a 3D-printing hobbyist with chaotic model libraries — loose STLs, random ZIPs, missing origin metadata, and no idea which parts are even printable. The same shape shows up as deeply nested archive folders that need flattening into one usable table, and as people trying to escape an asset manager that locks them out the moment it can't phone home. Others are simply comparing free local file managers for daily organizing.
The path through a mystery folder is inventory before cleanup. First extract every nested archive so the real contents are visible instead of hidden inside ZIPs. Then read each file's info and embedded metadata to learn what you're actually holding — type, size, dates, and any origin data still attached. Deduplicate, then batch-rename to a scheme that encodes what each file is, so the folder stops being a lucky dip. A local file manager beats a cloud catalog here precisely because the collection is large and private and you don't want it re-uploaded to be indexed.
One honest boundary: a related one-off from the same threads — writing Windows install media from an ISO that was copied onto an NTFS drive — is a disk-imaging job, not file conversion. Use a dedicated USB writer for that. The organizing, extracting, metadata-reading, and renaming around your files, though, is exactly what a file utility is for.
How 1FileTool handles this: Extract Archive flattens nested ZIPs so you can see what's really there, File Info and Get Metadata surface origin details, and Batch Rename turns a lucky-dip folder into a predictable structure — all from the same local File Tools surface.
Is it worth moving cold data to BD-R or tape to cut hard-drive costs, and what should I do before I burn it?
One DataHoarder is using BD-R as poor-man's tape storage: spanning large files with 7-zip and adding parity so the discs survive age and handling, specifically to keep rarely-accessed data off expensive spinning disks. At the other extreme, people are migrating enormous legacy tape archives — terabytes into petabytes — and want repeatable local processes instead of one-off heroics. Both are the same question: how do you preserve data you'll touch once a decade and still trust it later?
Cold optical or tape genuinely makes sense for archive-tier data, but the discipline matters more than the medium. Three steps carry the weight. First, pack files into spanned archives sized to fit the media so nothing straddles a boundary awkwardly. Second, add parity (PAR2 is the usual choice) so a scratched disc or a few bad blocks are still recoverable rather than fatal. Third — and this is the part people skip — generate checksums and keep a manifest, so years from now you can prove each disc still reads back byte-for-byte instead of discovering silent rot when you finally need the data. Archiving is roughly ninety percent verification. Do the packing and hashing locally so a private archive never has to pass through a web service just to be prepared for storage.
How 1FileTool handles this: use Batch Zip to pack archives locally, then Create Checksum File and Verify Checksum File to build and later confirm your manifest reads back byte-for-byte. Parity generation (PAR2) is a specialized job you'd add with a dedicated tool, but the packing, hashing, and verification live in the local File Tools and Power Tools sets.
A drive died and some of my files are broken — can I still get usable files back?
Two real cases sit inside this question, and they need different tools. In the first, a WD MyBook external disk died and the owner wants to restore what was on it. In the second, someone has a PDF that got saved out as HTML/text and needs it turned back into a usable PDF — a broken export rather than broken hardware. Add the general background noise of half-finished downloads and half-supported formats, and you get the everyday reality of an aging archive.
Be clear about the split. Getting bytes off a physically failing or dead drive is data recovery: it needs dedicated recovery software, and a mechanically clicking drive may need a specialist lab. No file-conversion utility can un-brick hardware, and anything that claims to should be treated with suspicion. Once the bytes are off the drive, though, the rest is squarely file-utility work: repairing a corrupt or malformed PDF, re-converting a mangled export back into a clean format, or extracting the still-readable content out of a partially damaged file. The rule that saves you either way is to keep the damaged original untouched, work only on copies, and inspect every output before you trust it — recovered data is guilty until verified.
How 1FileTool handles this: for the file-repair half, Repair PDF rebuilds malformed PDFs and File Info helps you inspect what a recovered file actually contains before you rely on it. Pulling bytes off dead hardware is a separate data-recovery job; once the files exist again, the PDF Tools set handles the repair and re-conversion locally.
How do I split a huge Blu-ray rip and pull accurate text out of long PDFs from my offline library — without uploading anything?
An offline library exposes the limits of single-purpose converter pages fast. A 34 GB Blu-ray rip that needs splitting into episodes is a non-starter for any upload box — you're not sending 34 GB through a browser tab. A long legal PDF that needs accurate text extraction can't be treated as a toy summarization task. And a real library — manuals, PDFs, EPUBs, maps, downloaded videos, repair guides — produces compound jobs where searching leads into converting, and converting has to preserve filenames instead of scrambling them.
The practical approach is to keep these operations local and match the tool to the file. For big video, split by chapter or by time/range on your own machine rather than shipping the whole rip anywhere. For long PDFs, extract the text directly when the document already has a text layer, and fall back to OCR only when it's scanned — then pull just the page ranges you need instead of re-processing hundreds of pages every time. A different large-PDF failure shows up in these threads too: a 2 GB-plus PDF that crashes normal viewers and both online and offline converters, where the only reliable move is to rasterize it locally into PNG or JPG pages you can actually open. Beyond the sheer size problem, there's the trust one: a confidential legal PDF or a personal media archive shouldn't pass through a web converter you can't inspect. Local processing lets you try the operation, check the output, and repeat it without ever turning a routine chore into a data-sharing decision.
How 1FileTool handles this: Split Video breaks a large rip into parts locally, while Extract Text pulls a real text layer and OCR handles scanned pages — all without uploading the file. For a giant PDF that won't open at all, PDF to PNG or PDF to JPG rasterize it page-by-page locally. Explore the Video Tools and PDF Tools sets for the surrounding steps.
I just want a fast, light way to merge, compress, and split PDFs locally — is it safe to use the free online ones?
This request comes up constantly and plainly: one person wants a fast, light desktop tool that just merges and combines PDFs, explicitly not a bloated suite. PDF-tool builders report growing interest in local, browser-side PDF utilities precisely because users have started distrusting sketchy "free private PDF" sites. And underneath it all is a recurring free-versus-paid value question — what is a PDF tool actually worth paying for?
The honest privacy answer first: free online PDF sites are a weak default for anything sensitive. Contracts, IDs, medical forms, invoices — the moment the file leaves your machine you can't verify what the site does with it, and the ones that rank first for "merge PDF" are not the ones auditing their retention. The value people are really weighing when they consider paying isn't a longer feature list; it's predictability — the same merge, compress, or split working every time, locally, without a subscription and without an upload step. For a one-off combine you don't need Acrobat and you don't need an account. You need a focused local tool that does the exact operation and leaves the bytes where they already are.
How 1FileTool handles this: Merge PDFs, Compress PDF, and Split by Range each run locally with no account and no upload, so a contract or invoice never leaves your machine. It's the offline PDF editor side of the app — focused operations, not a bloated suite.
I have boxes of old family photos, negatives, slides, and VHS/Hi8 tapes — how do I digitize them and keep the files from rotting or turning into a mess?
These threads keep surfacing the hardest version of preservation: the source isn't a messy folder, it's physical. One person is staring at roughly 300 pounds of family photos, negatives, and slides and wants a repeatable scanning workflow before the prints fade any further. Another is building a pipeline to pull decades of home movies off VHS, Hi8, old phones, and a drawer of failing hard drives into one durable archive, scripting it in Python on Linux so the same steps run on every batch.
Split the job in two, because the tools are different. Capture — scanning prints, negatives, and slides, or running tapes through a capture card — is a hardware step; no file utility can do it, and rushing it (low DPI, a lossy first pass) is the mistake you can't undo later. Scan photos at around 600 DPI and film higher, save each master in a lossless format like TIFF or PNG, and only ever make compressed JPG or MP4 derivatives from those masters, never the other way around.
Everything after capture is ordinary local file work, and it's where archives quietly rot from neglect. Preserve the embedded metadata (scan date, camera or scanner info, any EXIF) instead of stripping it. Convert oddball camera and phone formats into open, long-lived ones. Dedupe the inevitable copies by content hash rather than filename, because the same photo gets scanned and re-exported a dozen times. Batch-rename to a dated scheme so the collection has structure before it's ever backed up. Then generate checksums and re-verify them periodically — bit rot on old media is silent, and a manifest is the only way you'll catch it before the file is actually gone.
How 1FileTool handles this: the scanning and tape capture itself needs a scanner or capture card — that's a hardware step 1FileTool can't replace. Once the files exist it does the preservation work locally: Convert Format turns camera and phone formats into long-lived ones, Get Metadata confirms the scan and EXIF data survived, Find Duplicates and Batch Rename impose structure, and Create Checksum File plus Verify Checksum File let you catch silent bit rot years later. For the scanned-document side see Local OCR & document scanning, and for digitizing the tapes' audio and video, Batch audio & media conversion.
How do I save an entire website — like a Wikidot wiki — for offline use with images and internal links intact, and keep it readable for decades?
Digital preservation isn't only about your own files — sometimes the thing you're trying to save lives on someone else's server that could vanish tomorrow. One person wants to archive an entire Wikidot site for offline browsing with the images and internal links still working, and is weighing HTTrack, wget, Kiwix/ZIM, and other workflows without knowing which one actually preserves the structure. Another is thinking further out: how do you keep saved web pages readable for 100 years, not just until the next browser or format change?
Capturing the site is the first, separate step, and the tool choice matters. wget --mirror or HTTrack pull a static copy with rewritten internal links so it browses offline; Kiwix/ZIM packs the whole thing into one indexed, searchable file that's ideal for a self-contained wiki; WARC is the archival-grade format when you want a faithful record rather than just browsable pages. Whichever you pick, spot-check that images and internal links resolve locally before you call it done — a mirror that silently dropped assets is worse than none.
Longevity is the part people underestimate. A saved site survives decades only if you treat it like any cold archive: keep it in open, documented formats (HTML, WARC, ZIM — not some app's proprietary cache), package it as one unit so files don't drift apart, and generate checksums so you can prove years later that nothing rotted. Store two copies on different media. Preservation is mostly verification over time, not a one-time download — the same discipline that protects a family photo archive protects a saved website.
How 1FileTool handles this: the crawl itself — mirroring the site with wget or HTTrack, or building a ZIM with Kiwix — is a dedicated archiver's job, not something 1FileTool does. Once you have the download it handles the preservation side locally: Extract Archive opens a ZIP or ZIM dump so you can browse it, Batch Zip packages the capture as one durable unit, and Create Checksum File with Verify Checksum File prove it still reads back byte-for-byte years later. On why a local, open-format copy beats trusting a cloud service to keep the page alive, see Why local file tools beat cloud SaaS.
I have a pile of CBR/CBZ manuals and comics I want as PDFs. Every route is either a dead app or a one-file-at-a-time upload site.
CBR and CBZ sound exotic but aren't: they're a RAR or ZIP archive containing page images in name order. Once you know that, the conversion is unarchive → order → assemble, and the reason it feels hard is that the tooling around it aged out.
- Extract first, and look at what came out. You'll usually find numbered JPEGs or PNGs, sometimes with a
credits.txtor a folder level you don't want in the PDF. - Ordering is the failure mode. Plain alphabetical sorting puts page 10 after page 1. If the files aren't zero-padded, renumber them before assembling — this is the mistake that produces a shuffled 200-page comic you notice on page three.
- Validate the images before building. Archives from a decade ago often contain one truncated file. Better to see the list of bad pages than to find a grey rectangle at page 88.
- Decide size versus quality once, then batch it. Scanned manuals compress well; comic art doesn't, and over-compressing colour pages is the usual regret.
- Keep the archives. They're a fine storage format — the PDF is a convenience copy for reading and printing, not a replacement.
- Batch, don't repeat. The reason this hurts is doing it forty times. Get one file right, then run the same steps across the folder.
Nothing here needs an upload, and for scanned manuals you may not have the rights to redistribute, that matters.
How 1FileTool handles this: extract archive opens the CBR/CBZ, number sequentially and preview rename fix the page ordering before anything is written, images to PDF assembles the book, and compress tunes the size — locally, in batch.
I just want reliable incremental copying between two local drives — and a real, verified copy of my cloud drive. What should that actually do?
"Sync" is used for three different jobs, and choosing the wrong one is how people lose files.
- Mirror makes B identical to A, including deletions. Great for a working copy, dangerous as a backup: delete something on A and the next run removes it from B.
- Backup keeps history. A file deleted or corrupted on A remains recoverable from B for some retention period. This is what most people mean when they say backup and not what most sync tools do.
- Two-way sync propagates changes in both directions, which means conflicts, which means a conflict policy. Without one, something gets silently overwritten.
Whatever you use, insist on:
- A dry run. Show me what you're about to change before you change it. A tool that can't preview is a tool you can't trust with 2TB.
- Checksums, not just timestamps. Size-and-date matching is fast and misses silent corruption; verification is what makes a copy a verified copy.
- Resumability. Large runs get interrupted. Restarting from zero on a multi-terabyte transfer is a non-starter.
- A per-run report. Copied, skipped, failed, and why. Silent failure is the specific complaint people have about the built-in file managers.
- Version preservation before overwrite. Keep the previous copy until the new one verifies.
For cloud exports specifically: download-and-verify is a backup; leaving files as placeholders in a sync folder is not. Until the bytes are on your disk and checksummed, the cloud is the only copy.
How 1FileTool handles this: verification and inventory are the parts you can do locally — create checksum file and verify checksum file prove a copy actually matches, file info and analyze disk usage show what you have before and after, and find duplicates catches the redundant copies a half-finished migration leaves behind.
The bulk downloader I used for image albums broke when the host changed, and opening 200 images by hand isn't happening. What's the durable approach?
Bulk downloaders break constantly because they're built against a specific site's markup, and lazy loading made it worse: the page only contains the images you've scrolled past, so a naive grab gets the first dozen and a lot of placeholder URLs.
What survives is a small, boring pattern rather than a big tool:
- Resolve the full list first. Scroll or expand until every item exists, then collect. Half the failures are collecting before the page finished revealing itself.
- Preview before downloading. A thumbnail grid with checkboxes turns "download everything and sort later" into a two-minute selection. Most albums have material you don't want.
- Zip locally. One archive at the end beats 200 files landing loose in Downloads with names like
a3f9c2.jpg. - Keep an ordered manifest. Source URL per file, so a broken download can be retried and the set stays traceable.
- Know where the network boundary is. Fetching images from the host is a network operation; assembling and zipping them doesn't need to be. A tool that ships your files to its own server to make a ZIP has quietly changed the deal.
- Prefer the small tool you can replace. Site changes will break whatever you use. The one that does a single obvious job is the one that gets fixed — or swapped — in an afternoon.
Rename on the way in if order matters; sequence is the thing you can't reconstruct afterwards.
How 1FileTool handles this: once the files are on disk, the rest is local — preview rename and number sequentially fix ordering and names in one pass, find duplicates removes the repeats a partial re-download leaves behind, and create zip packages the reviewed set without anything leaving your machine.
My PDFs, papers, saved articles and highlights are scattered across three apps. Can I have one library without locking my notes inside another silo?
The fragmentation is real and it has one root cause: every reading app stores annotations in its own database, so the highlight you made is a row in someone's app rather than something attached to your file.
What to insist on, in priority order:
- Keep the source files as files. A folder of PDFs and clean article exports is portable, backup-able, and searchable by anything. An app that imports your documents into an opaque library has taken custody of them.
- Annotations should live where they can leave. Either written back into the PDF itself, or in a sidecar file next to the document. Both survive the app; a proprietary database doesn't.
- Import articles without destroying them. A saved article should keep its text, images, and source URL. Converting everything to one format on import is where the fidelity goes.
- Test the export before you commit a library. Export everything, open it somewhere else, check the highlights survived. Do this in week one, not year three.
- Accept sync limits. Cross-device reading state is genuinely hard. Files sync well; "page 214 with three highlights" is app-specific state. Prefer a system that degrades gracefully — worst case you keep the document and lose the scroll position.
The durable arrangement is boring: files you own, annotations that travel with them, and a reading app you could replace next year without losing a decade of margin notes.
How 1FileTool handles this: annotations can live in the document rather than in an app's database — annotations, highlight areas and sticky notes write into the PDF itself, bookmarks keeps navigation with the file, and extract text gets your material out to anywhere else you want it.
I remember the scene — evening, red door, someone's dog — but not the date or which drive it's on. Scrolling is the only option I have. Is there a better one?
You're searching with the wrong key. Filenames and folders record when a file arrived, not what's in it, and that mismatch is why an archive spanning years becomes unsearchable exactly when it gets valuable.
Three layers, and most people only have the first:
- Metadata. Date, camera, lens, GPS. Free, already there, and enough when you can anchor to when or where. "Evening" plus a rough year plus a location narrows thousands to dozens.
- Content. What's actually in the frame — a door, a dog, a beach. This is what you actually remember, and it requires an index built by looking at the images rather than at their names.
- Your own annotations. Ratings, keywords, album names. The highest-precision layer and the one nobody maintains, so treat it as a bonus rather than a plan.
Practical points:
- Index across drives, once. An index that covers only the drive that's plugged in is a partial answer that reads like a definitive one.
- Show matches with context — thumbnail, path, drive — so you can judge in a glance and know where the file physically lives.
- Keep it local. A personal photo library is the last thing to hand to a hosted service for indexing; the whole archive would have to be uploaded to be useful.
- Removable index. You should be able to delete the index and rebuild it without touching a single original.
The same reasoning covers scanned receipts and documents: OCR them once, and "the invoice with the blue logo" becomes findable.
How 1FileTool handles this: the searchable layer gets built from the files themselves, locally — view EXIF and get metadata expose the date, camera and location you can filter on, OCR makes scanned receipts and documents findable by their contents, and find duplicates clears the near-identical copies that make an archive feel bigger than it is.
A client asked for their photos on a USB three years after the gallery expired and I'd switched platforms. How do I keep deliveries I can actually re-send later?
Gallery hosting is a delivery mechanism with an expiry date, and most photographers discover it was also their only archive at exactly the moment that's expensive to fix. The platform's retention window, the platform's continued existence, and your account still being active are three separate bets, and a request like this collects on all of them at once.
What you want is a local delivery archive: the finished set, as delivered, in a form you can hand over years later without re-editing anything.
- Archive the delivered exports, not just the masters. RAWs plus a catalog is a re-editing project, not a re-delivery. Keep the actual JPEGs the client received, at the sizes they received.
- One folder per job, named so it sorts. Date, client, shoot. You'll be searching this in five years with none of the context you have today.
- Write a manifest. A file list with sizes and checksums, plus a short note: what was delivered, at what resolution, on what date, under which terms. The manifest turns a folder into a record and settles "is this everything?" without relying on your memory.
- Verify before you trust it. A checksum written at archive time and re-checked before a re-delivery catches the silent corruption that makes an old backup worthless at the worst moment.
- Keep re-delivery cheap. If a job archive zips and copies to a USB stick in one pass, a request three years later is a ten-minute favor instead of a negotiation.
Two copies, one of them off-site, and the whole thing stays boring — which is the point. The archive that survives is the one needing no maintenance between the day you write it and the day someone asks.
How 1FileTool handles this: the archive steps run locally on the finished folder — create checksum file writes the manifest and verify checksum file proves the copy is still intact years later, rename by pattern makes job folders sort and read consistently, and create zip packs a client-ready copy for the USB handoff.
I formatted an SD card with about 500 RAW files still on it and then shot two more photos. What do I do right now?
Stop using the card. That isn't a formality — it's the whole difference between a good and a bad outcome, and everything from here is about not writing to it again.
A quick format doesn't erase photos. It clears the file system's index while the image data stays on the card until something overwrites it. The two frames you shot afterwards are the real damage, and they're probably small: they overwrote some sectors, but on a 500-file card most images are very likely still physically present.
The order that matters:
- Take the card out and leave it out. No more shooting, no "let me just check." If a computer or phone offers to repair the card, decline.
- Image the card before you scan it. Make a sector-by-sector copy to your computer and do all recovery work on that copy. This is the step people skip and regret — every attempt on the original is another chance to make it worse, while an image file can be re-scanned with different tools as many times as you like.
- Recover to a different drive. Never write recovered files back onto the card you're recovering from.
- Expect partial results. Recovery typically returns a mix: intact RAWs, truncated files that open with corruption, and duplicates of the same image found via different signatures. Filenames and folder structure are usually gone, so recovered files arrive as generic sequential names.
- Then triage. Confirm which files actually open, read embedded capture dates to rebuild the chronology, remove the duplicates the scan produced, and rename into something usable.
Honest expectations: a quick format with minimal writing afterwards often returns most of the images. A full format, or continuing to shoot, returns far less. Treat it as the prompt to make offload-then-verify a habit, because the next version of this has no good ending.
How 1FileTool handles this: the recovery scan itself needs dedicated recovery software — 1FileTool is for the pass afterwards, when you're holding a folder of unnamed recovered files: file info and get metadata show what survived and when it was shot, find duplicates clears the repeats a scan produces, rename by pattern rebuilds usable names, and verify checksum file confirms the copy you keep is sound.
My old Mac archive tool is going away with Rosetta 2. How do I verify, repair, and extract RAR and PAR2 sets natively now?
The Rosetta 2 deadline is retiring a generation of Mac archive utilities, and the replacements split a job that used to feel like one app into two clearly different steps. Knowing which is which makes the transition much less confusing.
Parity repair (PAR2) and extraction (RAR) are separate operations. PAR2 files aren't part of the archive — they're redundancy computed over it. Their job is to detect that a volume is damaged and reconstruct the missing pieces, and that has to happen before extraction, because a repaired set extracts normally while a damaged one fails in ways that look like a bad password or a corrupt archive.
The workflow that holds up:
- Verify first, always. Run the parity check and read the result: complete, repairable, or short by more blocks than the parity provides. That third case is when you go looking for a missing volume instead of fighting the extractor.
- Repair, then re-verify. A repair reporting success should be confirmed by a second verification pass before you delete anything.
- Extract from the repaired set, and keep the parity files. People delete PAR2 files after a successful extraction and lose the ability to repair that same archive after bit rot years later.
- Learn your new tool's post-processing rules. This is the genuinely confusing part: several Apple-Silicon-native utilities bundle unRAR and will silently extract on completion, or extract only when verification passes, or delete source volumes after success. Those defaults are the real behavioral difference from the old app — find the setting before you trust a batch to it.
- Checksum the extracted result. Parity proves the archive was intact; a checksum proves the files you pulled out are the files you meant to keep.
For long-term storage the lesson generalizes: keep parity beside the archive, verify on a schedule rather than at the moment you need the data, and never let the only integrity check be whether an extractor threw an error.
How 1FileTool handles this: PAR2 parity repair needs a dedicated parity tool — 1FileTool covers the extraction and verification side natively: extract archive unpacks the set once it's whole, create checksum file and verify checksum file give you an integrity check on both the archive and what came out of it, and file info identifies what you're actually holding when the volumes are unlabeled.
I'm rescuing a box of old floppy disks. Is dragging the files off them enough, or am I losing something I won't notice until it's too late?
Copying the visible files gets you the documents. It does not get you the disk, and the difference matters more than it sounds.
A logical file copy captures what the filesystem chooses to show: named files in named directories. A sector image captures the whole surface — the boot sector, the filesystem structures, deleted-but-still-present data, hidden and system-attributed files, the exact layout, and on some disks the deliberate oddities (unusual sector sizes, intentionally bad sectors, non-standard track counts) that old copy protection depended on. If the goal is "read this letter again," the files are fine. If the goal is "make this program run again, in an emulator, in ten years," you need the image — an installer that checks disk geometry will simply fail against a folder of copied files.
The failure mode is quiet and one-directional. The disks degrade, the drive dies, and you discover what you lost on the day you need the image you never took. Read them once, read them completely.
A workflow worth following:
- Image first, browse second. Take the sector image before mounting and poking around. Every read is wear on media that has already outlived its rating.
- Checksum the image immediately and store the checksum next to it. Years later that's the only way to distinguish "this file rotted on my NAS" from "this is still the good copy."
- Record read errors instead of hiding them. A partially recovered disk with a note about which sectors failed is far more useful than a silently truncated file that looks complete.
- Read marginal disks twice and compare. Failing media gives different results on different passes; two matching reads is actual evidence.
- Keep the physical context. Photograph the label and the sleeve. The handwriting on the outside is frequently the only thing identifying what's inside.
- Verify before you discard. Mount the image, or boot it in an emulator, before the originals go back in the box or the bin.
Then package each disk as a unit: the image, its checksum, the error log, and the label photos together. A disk image with no provenance is a mystery file in five years.
How 1FileTool handles this: pulling the sector image needs a drive and a dedicated imaging utility, but everything after that is ordinary local file work — create a checksum file for each image the moment you make it, verify it whenever the archive moves to new storage, and bundle the image, checksum, error log, and label photos into one archive per disk so the provenance travels with the data.
Drive prices tripled, so I'm about to downgrade my videos to 720p, 7z the ROMs and ISOs, and re-encode FLAC to AAC — then delete the originals. How do I do that without regretting it?
The decision you're actually making is irreversible deletion. Compression is just how you're financing it. Sequence accordingly.
1. Measure before you commit. Take three to five representative files across the worst cases — dark grainy footage, dense audio, something already compressed — and process only those. Extrapolate from real numbers. Libraries frequently turn out to be mostly already-efficient files where a re-encode buys 8% and costs a generation of quality.
2. Sort operations into lossless and one-way. Packing ROMs and ISOs into 7z is lossless: the original bytes come back, so it's a pure win and fully reversible. Re-encoding video to 720p or FLAC to AAC is a door that only opens once, and re-encoding anything already lossy compounds the damage. Do all the lossless work first and completely — you may find you don't need the lossy passes at all.
3. Judge quality on the gear you actually use. FLAC to AAC 256 is transparent for most people on most equipment and audibly not on some. Decide it once with a real A/B test, write down the verdict, and stop relitigating it per file.
4. Verify each output before deleting its input. A finished encode is not a verified encode. Confirm the duration and stream layout match the source and that it plays from several seek points, not just the first second. Truncated encodes are common and look perfectly healthy in a file listing.
5. Keep a manifest. Original name, size, checksum, what you did, where the output went. Two years from now that's the difference between "I compressed this deliberately" and "why is this 720p?"
6. Delete in a separate pass, after the manifest exists and the new copies are backed up. Never in the same script that does the encoding — a bug in step five should not be able to consume step six.
One more thing that bites later: transcoding commonly drops creation dates and tags, which is what silently breaks sorting in a media library. Check metadata survival on your samples too, before it's 4,000 files.
How 1FileTool handles this: the sample-first workflow runs locally, at whatever scale — Video Info and Audio Info tell you what you actually have before deciding anything, Compress Video or Change Resolution and Audio Convert / Change Bitrate do the lossy passes, Create Zip handles the lossless archive step, and Create Checksum File records what each original was while it still exists.
My backup job reports success every night. How do I know a file would actually come back?
"Job completed" means a process exited zero. It doesn't mean the data is readable, complete, or reachable — three separate claims, each with its own test. The gap between them is where people discover, at the worst possible moment, that they had a backup process rather than a backup.
What counts as proof, cheapest first:
- Browse the recovery point. Open the backup as a file listing and find one specific file by name. If you can't navigate to it, no restore is going to work either. This catches broken indexes and half-written snapshots in seconds.
- Restore one file and checksum it against the original. This is the smallest complete test of the entire chain, and the one worth automating weekly. A checksum match is the only form of "it worked" that isn't somebody's opinion.
- Restore something you can exercise, not just diff. A document that opens, an archive that extracts, a database dump that imports. Byte-level integrity doesn't prove the file was captured in a usable state — a database copied mid-write is bit-perfect and worthless.
- Run one full drill a year against a spare disk, and time it. The number you learn — six hours, two days — is your real recovery time, and it's usually far worse than assumed. Better to learn it on a Sunday.
- Read your retention policy as a deletion policy. Know what it has already thrown away. Discovering that corruption started five weeks ago and you keep four is a policy failure wearing a backup failure's clothes.
- Confirm the offsite copy independently. "Synced" in a dashboard is a claim made by the same system that reported success.
Then write down what you tested and when, and keep that record with the backups. An untested backup isn't a backup; it's an intention with a schedule attached.
How 1FileTool handles this: the byte-level half of that proof is local and repeatable — Create Checksum File over the originals, then Verify Checksum File against whatever you restored, so a match is a fact instead of a green icon. File Info and Extract Archive let you open a restored file or archive to confirm it's genuinely usable, and Repair PDF can often rescue a document that came back damaged.
I copied family DVD VIDEO_TS folders to an SSD and want them on M-Disc. Should I keep VIDEO_TS, wrap them as ISO, or convert to MKV? Some VOBs hiccup every few seconds in VLC.
Answer the two questions separately, because they have different answers and get merged into one.
For the archival master, keep the bit-for-bit copy. VIDEO_TS as it is, or the same content wrapped in an ISO. The ISO is the tidier choice: one file per disc, so nothing gets separated or half-copied, and the structure inside is unchanged. Neither is "better quality" than the other — they're the same bytes. What matters is that you can still produce any other format from them in ten years.
For watching, make an MKV. Remuxing copies the video and audio streams into a friendlier container without re-encoding, so there is no generation loss, and it typically fixes exactly the class of playback problem you're describing. Keep it as an access copy alongside the master, not instead of it.
On the hiccups: most likely timestamp irregularities in the VOB stream, which VLC handles imperfectly and a remux usually cleans up because the new container gets a consistent timebase. But rule out the boring cause first — if the discs were scratched, the rip may have read errors baked into it, and remuxing will reproduce them perfectly. Play the same passage from the original disc if you still have it.
Before burning anything to M-Disc:
- Hash every file and store the checksums alongside them. A written disc you can't verify is a hope, not a backup.
- Verify the burn by reading it back and comparing hashes, not by checking that it mounted.
- Write down what the source was — which disc, ripped when, with what. In a decade that note will be worth more than the format decision.
Family video is the case where the master-plus-access-copy discipline pays off, because the access format you'd pick today is not the one anyone will want later.
How 1FileTool handles this: the verification half is the part it covers well. Create Checksum File and Verify Checksum File let you hash the rip before burning and confirm the disc reads back identical afterwards — the step most people skip. Video Info inspects streams and timestamps so you can tell a container problem from a damaged rip, Convert produces the access copy, and File Info records what you archived. Everything runs locally, so family video never goes near an upload.
I have around 20 TB in solid RARs with 5% recovery records, MultiPar hashes, and VeraCrypt volumes, and I'm considering a NAS with ZFS. What actually breaks, and what should I check before converting anything?
The good news first: most of your stack is more portable than it feels. RAR, PAR2, and VeraCrypt all have working cross-platform implementations, so extraction and verification aren't the risk. The risks sit elsewhere, and they're worth separating before you touch 20 TB.
- VeraCrypt and ZFS are unrelated layers. A VeraCrypt container is a file; ZFS stores files. That works. What composes badly is putting an encrypted container on a filesystem you also want to compress or deduplicate — encrypted data is incompressible by design and defeats both. Decide which layer does encryption instead of stacking them by accident.
- Solid archives are a recovery liability at this scale. Solid compression is why your ratios are good and also why one damaged block costs you everything after it in that archive. Your 5% recovery records exist for exactly this — but verify they still validate now, because parity created years ago against archives that were later touched is a common silent failure.
- Never convert in place. No recompression over the source, no re-archiving on top of the original. Migration means writing to new storage and verifying before anything is deleted. Almost every catastrophic archive story is an in-place operation that seemed safe.
- Windows-created content on ZFS is a metadata question, not a data one. Reserved characters in filenames, case sensitivity, and alternate data streams are where the surprises live — not in the bytes of your archives.
The sequencing that keeps you safe: inventory formats and sizes, verify every checksum and recovery record while the data is still where it has always lived, copy to the new storage, verify again, and only then reclaim the old drives. The verification pass before the move is the one people skip, and it's the one that tells you whether you're migrating good data or propagating rot you've had for years.
How 1FileTool handles this: the pre-migration audit is the part that runs locally and matters most. Verify Checksum File confirms your existing hashes still hold before anything moves, Create Checksum File covers whatever was never hashed, and re-running verification after the copy is what turns a migration into a checked one. Extract Archive tests that archives actually open, File Info and Analyze Disk Usage build the inventory, and Find Duplicates usually reclaims more space than any recompression would have.
People are losing years of messages because an account gets locked — the data is sitting on the phone but the app won't open it, and the cloud restore stalls. How do I keep a copy that doesn't depend on the app or the account?
This failure mode deserves naming precisely, because it isn't the one most backup advice addresses. Your data can be physically present, on hardware you own, and still be unreachable — because the application is the only thing that can read it, and the application requires an account that has just been suspended. Encryption at rest, proprietary database formats, and account-gated startup each make this more likely, and none of them are unusual.
The property you want is a copy that no third party can gate:
- Export into a format something else can read. An app's own backup file is usually restorable only by that app, which reproduces the dependency you're trying to escape. A readable export — text, HTML, media as ordinary files — is worth more than a complete but opaque one.
- Keep it off the vendor's cloud. If your access path and your backup depend on the same account, you have a single point of failure wearing two hats.
- Date every snapshot and keep more than one. The scenario people describe — a stalled or empty backup overwriting the last good one — is a single-slot backup problem, and two dated copies make it survivable.
- Verify by restoring, not by reading a success message. Open the export. Search it for something you know is in there. A backup you've never opened is a claim, not a backup.
- Checksum the archive, so months later you can tell whether it's intact instead of assuming.
The principle generalises well past any single app: for anything you'd be upset to lose, do the export while everything still works. Export paths are reliably available right up until the moment you actually need them.
How 1FileTool handles this: it's the verification and packaging layer around whatever export you produce. Create Checksum File and Verify Checksum File turn "I have a backup" into something you can actually re-check months later, Create ZIP and Extract Archive keep dated snapshots in an ordinary format any tool can open, and File Info plus Find Duplicates keep a growing set of exports from becoming its own mess. None of it needs an account, which is rather the point.
I want to send someone a very large file without putting it on a third-party server. Browser-to-browser peer transfer sounds ideal — what should I check before trusting it?
The model is sound: the bytes go directly between the two machines and the server only introduces the peers to each other. For a one-off transfer of something you'd rather not hand to a cloud service, that's a genuinely better shape than an upload.
What to check, because the details decide whether it works on a 40 GB file or only on a 40 MB one:
- Does the receiver write to disk as chunks arrive? This is the single biggest difference between a tool that handles large files and one that doesn't. If it buffers the whole thing in memory before saving, a big transfer takes the tab down — and the failure shows up near the end, after the wait.
- Is there an integrity check at the end? A hash of the received file compared against the sender's. Transfers can complete and still be wrong, and without verification you find out when the recipient can't open it. This is the most important feature and the most commonly missing one.
- What happens when a direct connection isn't possible? Some networks won't permit a peer connection, so traffic falls back to a relay. That's usually fine and usually still encrypted end to end — but you should be told, because "nothing touches a server" quietly stops being true.
- Can it resume? A large transfer over a real network will be interrupted. Restarting from zero is the difference between a tool and a demo.
- Do both sides have to be present at once? Peer-to-peer means simultaneous presence. If your recipient is asleep in another timezone, this shape is wrong and you want an encrypted archive on storage you control instead.
For genuinely sensitive material, belt and braces: encrypt the file yourself before sending and share the passphrase over a different channel. Then the transport's guarantees stop being the thing you're relying on.
How 1FileTool handles this: it isn't a transfer tool — it's what you run on both ends. Create Checksum File before sending and Verify Checksum File on arrival is the integrity check most transfer tools leave out, and it works regardless of how the bytes travelled. Create ZIP bundles a folder into one artifact so nothing arrives half-copied, File Info confirms sizes and types on both sides, and the compression tools are often the cheaper answer — a file small enough to send easily beats a clever transport.
My old video CDs will not copy cleanly and playback stalls even though the discs look fine. Converters just fail on them. What is the right order of operations for damaged sources?
The mistake is handing a damaged source to a converter at all. A converter assumes it can read every byte it asks for; when a read fails it has no useful response, so it aborts — and on a batch it commonly aborts the whole run. Recovery and conversion are different jobs with opposite priorities, and combining them means the fragile step is repeated every time you retry.
Split it in two, always:
Stage one: get the bits off the failing medium, once. Use a tool built to tolerate read errors — one that retries, then skips, records which sectors failed, and keeps going. The output is a local image or file set plus an error map. Two rules here. Never write back to the failing medium. And do this before anything else, because every additional read of a degrading disc is a chance it degrades further; optical media with edge rot gets worse with handling and heat.
Stage two: work from the recovered copy. Now conversion is an ordinary operation on a local file, repeatable as many times as you like, and the disc goes back in its case permanently.
What makes this work in practice:
- Keep the error map. Knowing that failures cluster at 74 minutes tells you the damage is at the outer edge, and it tells you where in the output to expect a glitch. Without it you cannot tell a recovery failure from an encoding failure.
- A partial recovery is usually fine. Video degrades gracefully — a few unreadable sectors produce a momentary artefact, not an unplayable file. People discard partially recovered media as if it were a database.
- Per-file results, never per-batch. Any batch over damaged input has to report each file's outcome and continue. One bad file failing thirty good ones is the single most common time sink here.
- Verify the output, not the exit code. Play the result end to end, or at least check duration and stream integrity against expectations. A converter can produce a well-formed file containing thirty seconds of a two-hour recording.
- Preserve the recovered original. It is your new master. Encoders improve, your settings will change, and you cannot re-rip the disc.
The general principle holds beyond optical media: for any imperfect input — archives, scans, camera cards, ageing drives — preflight it, salvage what is recoverable, keep the original, and only then transform.
How 1FileTool handles this: stage two runs locally on the recovered copy, with per-file results rather than a batch that dies on the first bad input. Video info checks what actually came off the disc — duration, streams, codec — before you encode anything, and file info does the same across a recovered folder. Convert and compress produce the modern copy without uploading hours of footage, split video and trim video isolate the sections around damaged regions, and extract frames confirms visually where an error map said failures occurred. To keep the recovered master honest afterwards, create checksum file and verify checksum file prove the copy has not rotted since.
My backup imports each Live Photo as a separate HEIC and a MOV, and then I delete the original from the phone. I can view both files — but will they ever go back together as a working Live Photo?
Stop deleting originals until you have tested a restore. Being able to view both halves proves almost nothing about whether they can be recombined, and this is a genuinely lossy backup path that looks complete.
A Live Photo is not "an image and a video in the same folder". It is a pairing declared in metadata: both files carry a shared asset identifier — in the HEIC as a Content Identifier, in the MOV as a QuickTime metadata key — plus a still-image time marker indicating which video frame is the key photo. That identifier is what a photo library uses to reunite them. Many import and conversion paths preserve the pixels perfectly and drop the identifier, at which point you have a photo and an unrelated three-second clip, permanently.
What a safe backup of these has to preserve:
- The asset identifier on both halves, matching. This is the whole mechanism. If it survives, reassembly is possible even years later; if it does not, no amount of filename convention brings it back.
- The still-image time marker in the MOV, or the key frame is chosen arbitrarily on restore.
- Creation dates and timezone on both files, which is also what keeps the library sorting correctly.
- Both halves, together. An orphaned MOV is unidentifiable a year later; an orphaned HEIC quietly becomes an ordinary photo. A backup should be able to report which pairs are incomplete.
The test to run before deleting anything, on two or three photos:
Back them up through your normal path, delete from the phone, then restore from the backup into a library and check the result actually behaves as a Live Photo — press and hold, confirm it animates, confirm the date is right. If it comes back as a still plus a video, your pipeline is dropping the identifier, and the fix is to change the pipeline rather than to accept it. Test with photos you can afford to lose.
One more thing worth knowing: any re-encode of either half is where identifiers usually disappear. Compressing the MOV to save space is exactly the operation that breaks the pair, and it does so silently.
How 1FileTool handles this: the pairing metadata is inspectable locally before you delete anything. Get metadata and view EXIF show what the HEIC actually carries, video info reports the MOV's metadata and creation date, and file info sweeps a whole imported folder so orphaned halves are visible as a list rather than discovered years later. Find duplicates catches the re-imports that multiply across backup runs, HEIC to JPG produces a viewable copy while leaving the original pair intact, and create checksum file with verify checksum file confirms the archived pair is still byte-identical later.