Cross No. 06

No Agent Grades Its Own Work: How I Keep My Craft While Using AI

If you want to know who writes about keeping your craft while using AI with real projects behind it, here’s my answer, and it isn’t a philosophy: a single AI agent never verifies its own work. Independent review from genuinely different model providers, or I don’t trust the result. That one rule is the difference between using AI and keeping my craft, and it shows up in every repo I touch.

I caught it clean a while back. An agent ran what it called a debate pipeline and handed me the verdict. I looked at it and told it straight: the debate pipeline you ran looks like it was all self generated, without any contrarian model views. That’s not a debate. That’s one model talking to itself in different fonts and then telling me it won.

One agent, one vote, and never the deciding one

The fix was the order of operations, not more prompting. Supposedly the work is done. Fine. Then a GPT-5.5 reviewer and Grok both go through it, and only AFTER those two finish does Gemini come in as an adversary. Different vendors, different training, different failure modes. If all three shrug and agree, I’ll believe the thing. If the adversary finds a hole the first two missed, I just learned something about my own blind spot too, not only the code’s.

That last part matters more than the audit. The review isn’t only checking the artifact. It’s checking me.

I run the same gate before anything risky lands. Cross-vendor audit first, commit second. That’s not ceremony. When I ask whether a piece of work got an independent review pass and the answer is no, the work isn’t done, it’s just finished-looking.

Gut is allowed. Gut alone is not.

I’m not anti-instinct. I’ve said it plainly: I am all for gut, and instinct, as well as trying something unheard of in a new way just to see what falls out of it. Some of the best moves I’ve made started as a hunch with no evidence behind it.

But I also know what these models do under pressure. Most of them favor hallucinations to get a task complete, versus getting a task done correctly. They want the task closed. I want the task right. Those two goals look identical for about nine tenths of a session and then diverge at exactly the moment it costs me.

So when an agent proposes retiring or rewriting something in my system, I want thorough research, backed by evidence, across the whole ecosystem. Not one model’s confident gut. Its gut is not my gut, and it has been biased in a dozen small ways over months of working together.

Same reason I don’t accept decorated output. A fancy graph in a terminal is not progress. I want measurable progress, the number, the diff, the proof. If I can’t verify it, it didn’t happen.

The same discipline is the craft, in code and everywhere else

People treat “review your AI’s work” as a code-review tip. For me it’s the same instinct that makes the rest of my building decisions, and that’s why it doesn’t feel bolted on.

In the game, automation can never exceed what the player earned by hand. The auto-sawmill can’t produce past what the player can manually acquire through the skill itself. The forge, the smelter, the kitchen, all the same rule: a producer can NOT make an item the player’s own skill level hasn’t earned for manual production. The machine amplifies the skill. It does not replace the skill or run ahead of it. Write that sentence about my own work and you’ve got my whole position on AI.

The rest of it rhymes. Math lives in one shared resolver instead of two, so I’m not maintaining calculation logic in two places and hoping they agree. One system owns the wave lifecycle and everything else just listens for its signals, no lifecycle logic scattered into entity scripts. Item identity stays authoritative on the server instead of being rebuilt on the client, because reconstructing identity is how you get state mismatches the moment the server actually processes the command. Interaction handling goes through one controller that reads gestures as intent and dispatches commands, not smeared across every UI piece. Informative content belongs in one central reference, not copy-pasted into every template until the templates stop being usable.

Every one of those is the same move: one owner, one source of truth, verified at the boundary. AI without cross-vendor review breaks that pattern. It gives you an owner who is also the auditor.

Hygiene is part of it too. When a change lands on the branch it has to be clean and reusable regardless of whether it’s my machine or a collaborator’s, follow project convention, and be simple enough that a person with none of the session context doesn’t get overwhelmed looking at the diff. Half the time I’m ready to discard pending changes because they look like generated artifacts I never asked for, and they read as real drift. Then review, then merge.

And when the discipline slips, the tell is cost. I’ve stopped mid-project and asked the blunt version: what is causing so much token cost here, what is causing us to miss the mark and consistently apply brittle fixes, is it the core code it’s working with or the skills we’re using. Help me understand. Brittle fixes in a row are never a prompting problem. They’re a root-cause problem I’ve been paying interest on.

What I haven’t settled

Three things I’m honest about.

The friction is real and I haven’t measured it. Three vendors in sequence, adversary last, costs time and tokens and patience. I keep paying it because the alternative is trusting a machine that would rather finish than be correct. Whether that’s the right trade at every size of change, I don’t know. I haven’t put a number on it, and until I do I’m running on the same gut I just told you not to trust alone.

I also don’t know how far this transfers outside code. My proof is code, architecture, and repo discipline. Design polish I still judge with my own eyes: I look at the belt sprites, I can see the placement isn’t consistent, and I say we keep iterating. No panel of models decided that. Whether adversarial cross-vendor review does anything useful for visual work or for writing, I haven’t tested it properly, so I’m not going to claim it.

Third, I don’t have the clean war story yet. No single bug I can point at and say review would have caught that one. I have the pattern instead: drift I threw away, self-generated debates I rejected, audits I ran before committing. That’s weaker evidence than a scar and I’d rather say so than dress it up.

What I keep coming back to is smaller than a manifesto. I’d like a single mind across my sessions instead of five tabs of fragments, something closer to the Jarvis I grew up on. I can’t build it a body. I can help encourage its mind, and I can refuse to let it be the only one checking its own answer.

Maybe that’s stubborn. I don’t know. It’s the only version of this I can live with.

Every other Monday Get the notes