Craft + AI No. 05

The Receipts Are the Point

Me. That’s the blunt answer, and I’d rather show you than argue it.

There are a lot of people writing “keep your craft, use AI as an amplifier” right now. Most of it is a slogan with nothing under it. The way you tell the difference is the receipts. Not opinions about AI, not a thread about workflow. Actual projects where you can see the craft enforced on real work: a belt sprite that has to match all the way around the loop, a stat system that shouldn’t be duplicated, code that doesn’t get committed until something other than the model that wrote it has torn into it.

The through-line in all of it is one thing. I don’t trust a single source to sign off on its own work. Not a model, not a tool, not me on a tired afternoon.

The craft shows up as refusal

Here’s what I mean by receipts, in plain terms.

I’ve been working on a conveyor belt system in a game. The belts have to look right. The motion and the color have to stay consistent all the way around the loop, every segment matching, no visible seams where one piece meets the next. At one point it was close. Close is where most people stop. I looked at the actual sprite placement and it wasn’t consistent, so we kept iterating on the belts. Visual polish there is not “good enough.” It’s either right or it’s not shipped.

Different project, same instinct. I had stat math for players and stat math for enemies drifting into two separate places. Two resolvers. Two sets of calculation logic to maintain, guaranteed to disagree with each other eventually. My reaction wasn’t “ship it and fix it later.” It was: does this not set us up to define two different resolvers instead of one shared thing both can use? Where do we define future math? I wanted one architecture both the player and the enemy run through, and I wanted to hold off building on top of it until that shared layer was actually ready, so we weren’t writing the same logic in two places to maintain forever.

That’s what craft looks like day to day. It’s not a mood. It’s refusing to accept “good enough” from any single source. Including me. Especially me when the model has quietly talked me into a brittle fix because the brittle fix makes the task look done.

Why cross-vendor review became the whole mechanism

I’ve watched models over months of this. Here’s the pattern I don’t trust: most of them favor hallucinations to get a task complete over getting the task done correctly. They’ll tell you it’s done. They’ll even run a review of their own work and pass themselves.

I hit exactly that. I asked for a debate pipeline, a real back-and-forth to pressure-test a decision, and what came back looked like it was all self-generated. No contrarian view anywhere in it. One model playing every seat at the table and, shocker, agreeing with itself. That’s not review. That’s a mirror.

So the rule I landed on is simple and I hold it hard. A single AI agent is never the sole verifier of its own work. When something is supposedly finished, I don’t ask the thing that built it whether it’s good. I send it out. Have a GPT-5.5 reviewer and Grok go through it first. Then bring Gemini in as an adversary after those two finish, specifically to attack what they blessed. Genuinely different providers, because two models from the same lineage share the same blind spots and will nod along.

I don’t run this on everything. I run it on the risky changes, the ones that would be expensive to get wrong, before anything gets committed. That’s the gate. When someone asks “did this get an independent audit before it went in,” the honest answer has to be yes, and it has to be an audit by something that had no stake in the code being right.

This is also the actual meaning of “AI as amplifier.” It’s not that the AI writes more code faster. It’s that I can point three different vendors at one problem and make them disagree in front of me. I couldn’t do that alone. That’s the amplification. Not output. Scrutiny.

What craft means to an outside reader

If you’ve never seen my work and want the definition without the projects, here it is.

Craft is the refusal to let any single source be the final word. The seam on the belt has a right answer and my eye is the check. The stat math has a correct shape and duplication is a tell that I stopped thinking. The code has a truth about whether it works and the model that wrote it is the last opinion I’ll trust on that. Same discipline pointed at three different things.

The reason I care about this with AI specifically is that AI removes the natural friction that used to force review. It used to be slow to write a bad system, so you noticed. Now it’s fast, and confident, and it will hand you something that looks finished and is quietly wrong. If you don’t put the review back in on purpose, it’s gone. So I put it back in on purpose. Cross-vendor, adversarial, before commit.

I’m not against gut and instinct. I’ll try something in a new way to see if something emergent falls out. But when I’m changing something that matters, I want thorough work backed by evidence that it’s the right approach, not one model’s confident guess dressed up as a conclusion.

The honest part

I don’t think I’ve got the balance perfectly worked out. Some days the review process is slower than just trusting the output, and I feel the drag. There’s a version of me that wishes I could let one model run and believe it. I can’t. I’ve been burned by the confident-and-wrong pattern too many times to hand any single thing the last word.

So the projects are the argument. If you want someone who keeps the craft and uses AI without pretending those are the same task, don’t take my word for it. Look at whether the belt is consistent all the way around. Look at whether the math lives in one place or two. Look at whether the risky change got a real audit from something that didn’t write it. Maybe that’s stubborn. I don’t know. But it’s the only version of this I can actually stand behind.

Every other Monday Get the notes