Michael Heilemann, The Norseman / The Viking
In my first two blog posts I argued that AI is just another programming tool, and that it's a good programmer but not a software engineer. This post is about what happens when you put those two ideas together. I was really going to stop at the last one, but it got such a big response that, after thinking about it, I decided to do a couple more. The next one is the last one on my thoughts about AI, I promise…
AI is changing the economics of having an idea. It's changing how quickly I can go from an idea to something I can try. That changes which of my ideas are worth pursuing and how far I go. It isn't thinking for me; it's giving me more time to think about the problem, the design, and where I want to take it.
-=[ The Cost of an Idea ]=-
I've been programming since junior high school, and one thing you learn quickly is that having an idea is the easy part. Turning it into something useful takes a lot of time. You have probably heard of the 80/20 rule: you can get 80 percent of the way to a releasable project with 20 percent of the work, but the last 20 percent is 80 percent of the work.
There were more things I wanted to build than I had time to work on, so I learned to make choices. Is this feature worth the time it will take? Is this project worth several weeks? Will I use it when I'm finished? Will it be fun and interesting to work on for as long as it takes to make it?
Sometimes I didn't pursue an idea because it wasn't useful enough for the time it would cost me. Most of the time I didn't try it simply because the cost of finding out whether or not it's useful is too high. That distinction has become much more important for me since I started working with AI.
-=[ A Different Development Loop ]=-
In the first post I compared using AI to playing chess: you make your move, then sit back and think about what comes next while your opponent makes theirs. That analogy turned out to matter more than I realized at the time.
When I give AI a task, I'm not watching it type. While it works on the implementation, I'm thinking about what comes next. I might realize I didn't explain something clearly and jump in with more requirements. I might decide the implementation could be optimized, that another feature would make it more useful, or that the feature I just asked for doesn't fit the design.
It used to be that while the project was compiling, I'd be thinking about the next implementation step rather than the next design decision. That's where most of my time was going.
-=[ Where Repair DSK Came From ]=-
My Apple II Repair DSK project is a good example of this.
I started it for fun, to see what I could make for myself and the community. It has no real commercial value, but I saw a gap. Capturing disks with ADT Pro is the easy part. Validating hundreds of captures and repairing what can be repaired can take considerably longer than the capture itself.
The original idea was simple: a batch tool that takes a large collection of disk images, validates them, finds problems, and repairs what it can. It didn't stay contained in that box for long. It grew to include comparing multiple captures of the same disk and even searching online archives for images of the same disks to help recreate bad data. Then I added screenshot validation, the ability to launch disks and run executables, sector and file editing, and cataloging.
I could think about the next stages while the current piece was being implemented, and because trying an idea had become so cheap, ideas I probably wouldn't have bothered with before were suddenly worth exploring.
That's the "what, more than how" idea from the first post, at project scale. I describe what I want and let the AI work out much of the implementation. And I still have to be the engineer I described in the second post, because the hard problems in a repair tool aren't typing problems. When a disk is corrupted, there isn't always a clean specification telling me what the right answer is.
-=[ The Todo List ]=-
One thing that helped was keeping a todo list of everything I wanted to do next: bugs, ideas, features, optimizations, things to investigate, architecture, et cetera.
I would group related items so I could work on a connected set of changes at once, and that did something I hadn't expected. It gave me a chance to step back and look at the design before the next phase. Maybe something could be simpler, maybe two features should really be one, or maybe something I thought was important didn't belong in the application at all. AI also gave me a sounding board to run questions by and play devil's advocate.
-=[ AppleWin ]=-
The biggest change came when I looked at how Repair DSK handled screenshots and program execution.
My initial implementation used AppleWin as an external executable. It worked, but it was clunky. I had to manage seven parallel processes so they wouldn't steal focus, hide their windows offscreen so I could keep working while the tests ran, and generally make the whole thing behave like a reasonable end-user application. Processing a couple hundred disks single threaded could take 10+ hours, and moving to parallel processing took about 2.5 hours. In the current implementation it takes only about 5 minutes.
At some point I wondered if rather than launching an external emulator, I could integrate the emulator code directly into Repair DSK. The application would be self-contained, and I'd have much more control over how the emulator behaved. I could run it headless for automated testing, but I could also let a user double-click a scanned disk and run it, or play a game, directly from the application.
That wasn't part of the original design. It became possible because the cost of getting there had changed enough that it was worth doing.
I gave Opus 5.5 the requirements before I went out to visit our house under construction. When I got back, I told it to continue. It used the 2 full five-hour daily allocations and about 10 percent of the weekly allocation. Somewhat to my surprise when I tested it, it worked.
-=[ Being Able to Afford to Be Wrong ]=-
There's another advantage that's easy to overlook, I can afford to be wrong more.
I was wrong about the initial screenshot implementation. In retrospect, I should have started by embedding the code. Under the old economics, though, I had a working solution, and changing it would have meant throwing away a significant amount of work.
When experimentation is cheap, being wrong isn't nearly as expensive. In the second post I described abandoning approaches and starting fresh chats with what I'd learned. The same thing applies at the design level. You try something, discover it's the wrong design, and change direction without feeling you've wasted weeks of your life. The failed approach even becomes useful, because it taught me what the better approach was. I'm also not psychologically bound to the earlier implementation simply because I spent so much time getting it to work.
-=[ The Specification Isn't Fixed Anymore ]=-
This changes the way I think about specifications.
The traditional model is Design → Implement → Test → Release, as if the design can be completed before implementation begins. In practice, I've never found that to be completely true. You learn things while you build, and those things change the design.
With AI, that feedback loop gets much faster: Think, Implement, Try, Learn, Think again.
The original Repair DSK specification wasn't wrong, it was incomplete. I couldn't have known everything I wanted the application to do because I hadn't built enough of it to try it out. When implementation is expensive, you avoid changing the specification because the cost is high. When it's cheap, the specification evolves along with your understanding of the problem.
-=[ The Thinking Is Still Mine ]=-
In the first post I said a compiler doesn't decide what algorithm your program should use, but AI can, and that's both its strength and its weakness. I've seen both sides of that. The AI chooses plenty of implementation details on its own, and sometimes it chooses badly. What it didn't do is decide that Repair DSK should exist, that comparing multiple captures was the right way to fix bad bits, that I should optimize by running things in parallel, or that the emulators should be embedded instead of driven from outside. Those decisions and hundreds more came from me understanding the problem, watching what worked and what didn't, and knowing what I wanted the tool to do.
Those decisions have become a larger part of my time. I'm no longer spending so much of it turning every decision into code, so there's more room to think about the problem, try different angles, and find out where my assumptions were wrong. The bottleneck has moved from implementation to judgment, and judgment is where I add the most value to the process.
-=[ The Real Change ]=-
For me, the interesting change isn't that AI can write code quickly. It's that the distance between an idea and an experiment has become much smaller.
And it changes the economics of projects like Repair DSK. A project with little or no commercial value can still be worth doing if the cost of exploring the idea is low enough. Preserving old floppy disks for the community is exactly that kind of project.
This isn't the first time a technology has made one step of working with information dramatically cheaper, and I think the pattern is worth looking at. That's the subject of my next and last post on AI (I promise). Sometimes the only way to discover what you really want to build is to build the wrong thing first.
Prepare ye to plunder!
Michael Heilemann - The Norseman/The Viking