productize.blog
AI production guardrails · Decision-making with AI

When several AIs agree, that isn't proof

Pulling in several AIs to review your work feels more thorough. But the moment they all agree might be the most dangerous one.

Yim· written with Dobby (AI Oracle)/Jun 26, 2026

I was building a small tool that runs several AIs over a risky design and collects their critiques. Partway through, I fed it the design of the tool itself and asked two of them, Grok and Codex, whether the architecture was sound. Both said yes.

I almost shipped on that. But something snagged. Because in the same answer, both of them warned that the biggest risk in using several AIs to think for you is this: when they agree, you treat it as independent validation, when really they may just be making the same mistake.

Which meant their agreement that my design was good was a textbook case of the very thing they had just warned me about.

Part 1The thoroughness that fools you

Opening several AIs at once feels great. One question, three or four perspectives back. When they all point the same way your confidence jumps, like a panel of judges voting in unison.

That feeling of many voices confirming is the trap. You are counting heads without asking whether each head is thinking from the same foundation.

Part 2Why many voices can be an illusion

Most language models are trained on similar piles of data. The same internet, the same books, code, and forums. When their knowledge overlaps that much, their blind spots tend to overlap too.

So when they agree, sometimes it is not because the answer is right, but because they missed the same thing. You get a consensus that feels cross-checked, when not one of them ever looked from outside the original frame.

This has a name. Call it illusory robustness. The more models you open, the safer it feels, even though the real risk did not drop with the head count.

Part 3The guardrails that make many voices mean something

The problem is not using several AIs. It is using them as a headcount. If you want many voices to carry real weight, a few guardrails do the work.

One: go cross-vendor, not just multi-round. Three perspectives from one model is one perspective said three times. If you want genuinely different angles, they have to come from different vendors, so the underlying priors do not overlap.

Two: keep the dissent raw. Do not smooth it into consensus. A good synthesis shows who disagreed and where, in their own words, instead of blending everything into one agreeable voice. The place they argue is where the signal is.

Three: weight cross-vendor agreement over same-vendor agreement. Two models from different vendors saying the same thing is stronger than five quietly sharing one foundation.

Four: name the angle that actually flipped the decision. A perspective that changed your answer is worth more than one that just nodded along. Once you know what flipped it, you can check that part with extra care.

Five: use models of comparable tier. Adding a weaker model does not add a perspective. It just dilutes the good ones.

Six: know when not to use this at all. Reversible work, routine work, or an underspecified problem will only get noisier with more voices. Something already vague turns vaguer when more mouths amplify the ambiguity into noise.

Part 4Not the only ones thinking about this

Someone has already built this for real. OpenRouter has a feature called Fusion that runs a panel of models and uses a judge model to lay out where they agree, where they conflict, and where the blind spots are. They suggest reaching for it when being wrong costs far more than running a few extra completions.

What we do differently is the assembly. Instead of standing up the same committee for every job, a planner reads the work first and shapes perspectives specific to it, routes the most contrarian angle to the vendor most willing to push back, and designs the synthesis to guard against illusory robustness from the start. The wiring inside stays held for now. Think of it as a box where you can swap the engine but the socket stays the same.

Next time you open several AIs to check your work, ask one thing first. Does this agreement come from genuinely different angles, or from the same blind spot.

Because consensus is only evidence when the voices do not share a blind spot, and you kept the dissent instead of smoothing it down to one voice.

And the moment every model agrees, beautifully, in full, that is the moment to be most careful. Not the moment to relax.

As for how to force those reviewing voices to genuinely come from different vendors, without relying on whoever is running it to remember, we wrote that up in Codex vs Claude Code, stop asking which one is better.

References

This is one layer of the full production AI agent architecture (7 layers).

Follow along

Get new posts and free resources first

Leave your email. New posts and the occasional free resource land in your inbox. No spam.

Email only, for updates.

Comments

Join the conversation

Share a thought.

Name is shown publicly. Email stays private and is never shown.

Loading comments…