/> WELCOME TO THE NORSEMAN'S CABIN </

═══════════════════════════════════════
Archiving, Validating and Repairing Apple ][ Disks

Michael Heilemann, The Norseman / The Viking — Field notes from the disk revival project: Apple II archiving

-=[ Every One of These Disks Is Fading ]=-

If you've ever tried to digitize an old Apple ][ floppy disk collection, finding the disks turns out to be the easy part. Trusting them is the hard part. A 40 year old 5.25″ floppy can look perfectly fine and still be wrong: a flaky drive head, a marginal read, a sector that comes back as garbage without ever throwing an error. ADT Pro will happily hand you a .dsk file built out of that bad data, and you won't know it until the day you actually try to use it and get far enough in to hit the damaged data.

I recently picked up a broken Unenhanced Apple ][e to restore and pass along. It came with 173 disks, most of which have data on both sides. Once I started looking at the machine itself, the problem list grew fast: a bad power supply, a drive that needed recalibrating, and a second drive that's going to need repair. It's a Rana Elite 3, and if anyone reading this can help me fix the slow drive speed issue, I'd love to hear from you. I had one of these 40 years ago storing BBS data, and it burned in the Palisades fire, so getting this one running again means more to me than just fixing a random drive.

Before any of these 173 disks (302 sides) go up anywhere (I plan to create a validated .dsk archive on the site here for download), every single image needs to be verified. Doing that by hand would take forever, and even then, I couldn't be sure I'd caught everything short of playing every game to completion (I still won't be able to do that with automation). So I built a verification tool instead and leaned on AI to help me get it built faster than I could have working by myself.

What started out as a simple batch disk verification tool quickly turned into something more: a validation and repair tool, with MAME and AppleWin screen capture, sector editor, file editor and launcher among other things. There are a lot of these disk images already floating around online, and it occurred to me that instead of treating every bad read as a dead end, I would try to cross-check against them.

-=[ The Golden Rule ]=-

One rule guided the design from the start: nothing that gets read off a disk after the transfer ever gets modified. Whatever lands in that first directory stays exactly as it arrived. That means the whole pipeline can be rerun on the same source data at any time. If something changes downstream, I can delete the working directories, merged disks (merged from multiple scans), repaired disks, and downloaded reference disks, and regenerate everything from the untouched originals. More on how that works below and detailed in future blog posts that I hope to write and release once per week.

-=[ Why "It Copied Fine" Isn't Good Enough ]=-

A successful transfer only tells you that the bytes made it into the image. It doesn't tell you that they are correct. You need to understand the filesystem supposedly living inside those bytes: DOS 3.3's VTOC and catalog, ProDOS's directory tree. Does a catalog entry's track/sector chain terminate? Does it point anywhere off the edge of the disk? Does the disk's own bookkeeping agree with what's actually sitting in those sectors? A disk can pass a checksum and still fail every one of those tests.

Once the physical disk itself degrades past the point of a clean read, and 40 year old floppies do that on their own timeline, whatever's in the archive is what that title is, permanently. A silently corrupted disk image that nobody catches now can become the canonical, wrong copy of that game.

-=[ Why Did I Even Do This? ]=-

Mostly just to see if I could. Nobody asked for a verification and repair pipeline for a stack of 40 year old floppies. There's no deadline, no client, no one waiting on it. Somewhere between parsing the VTOC and writing the bit map check, "can I build this" turned into "I want to build this right." I hope at least some of you find it useful.

-=[ What This Series Covers ]=-

This is the first post in a 12-part series exploring the real pipeline built in layers, each layer earning its place by catching something the layer before it didnt cover. Here's what's coming:

Part 1 (You are here): Every One of These Disks Is Fading.

Part 2: Talking to the Apple II: seriALL, a Floppy Emu, and Two Drives

Part 3: The Calibrated Capture: Why Every Disk Gets Read More Than Once

Part 4: Reading DOS 3.3 the Way DOS Does

Part 5: The Bit Map Doesn't Lie

Part 6: ProDOS, Briefly.

Part 7: Trust No Single Read: Consensus Repair

Part 8: Borrowing Bytes: Cross-Referencing Against the Archive

Part 9: Protected Disks, Honest Limits, and Shipping It

Part 10: Booting Blind: Automating MAME and AppleWin to Catalog What Actually Runs

Part 11: Watching the Screen: OCR, Best-Frame Selection, and Catching a Bad Boot

Part 12: Inside the Bytes: The Sector Editor and File Editor

-=[ Applying Archivist Principles ]=-

None of this is really new, the same basic ideas show up in other forms of digital preservation: make multiple reads, compare them, cross-check against known-good copies, and don't put too much faith in a single result. The real challenge is doing that consistently across an Apple II disk collection, rather than imaging a disk once, checking that it looks okay, and moving on.

Next week: the hardware. What's plugged into what, why a Floppy Emu is standing in for a boot disk, and why two drives instead of one turned out to matter more than I expected.

Prepare ye to plunder,
Michael Heilemann - The Norseman/The Viking

← ENTRY 2 ENTRY 4 →