← All posts
·6 min read

Why AI Can't Judge Its Own Design Output

Ask an AI builder to critique the interface it just generated and it will approve. That is not a bug you can prompt around — it is a structural blind spot, and understanding it changes how you ship.

Klyxx TeamKlyxx Team

Try this experiment. Generate an app with any AI builder, then paste a screenshot back into the same model and ask: "Is this good design?"

It will say yes. Maybe with a caveat or two — "consider adding more whitespace" — but fundamentally, yes. The model that produced a generic interface will look at that generic interface and see good design.

This isn't the model being polite. It's something more structural, and once you see it, it explains a lot about why vibe-coded apps plateau at "fine" — and what it actually takes to get past that.

Where a model's taste comes from

An AI model's sense of what good design looks like is learned from its training data: millions of interfaces, design systems, component libraries, and landing pages scraped from the web. Its "taste" is, in a very real sense, the statistical average of everything it has seen.

That average is not bad. The web's design consensus — legible type, sensible spacing, familiar patterns — is encoded in there, and that's why AI builders produce usable interfaces on the first try. Ten years ago, getting a non-designer to a usable interface took weeks. Now it takes a sentence.

But the same mechanism that makes the output competent makes it uniform. When the model generates a design, it reaches for the center of the distribution. And when the model evaluates a design, it measures against... the center of the same distribution.

Generator and judge share one set of eyes.

The blind spot, precisely

Here's the uncomfortable logic. The most common failure mode of AI-generated design is that it looks like AI-generated design: the same gradients, the same floating cards, the same three timid font sizes. Call it the template look.

Now ask what the model would need in order to detect the template look. It would need a reference point outside the template — some notion of "distinctive" that isn't just another sample from the same distribution. It doesn't have one. The template look isn't a deviation the model can flag; it's the model's null hypothesis. It is what "correct" looks like from the inside.

So when you ask the generator to critique its own output, you're asking the average to notice that something is average. The question doesn't parse. The model checks the design against its learned consensus, the design is the learned consensus, and the verdict comes back: looks good.

This is why "make it more unique" prompts disappoint. You're not adding new information — you're re-rolling dice inside the same distribution and hoping for an outlier.

"Just use a different model" doesn't fix it

The obvious workaround is to have one AI generate and a different AI judge. It helps less than you'd expect, because the major models are trained on heavily overlapping data. They share most of the same monoculture. A second model will catch genuine errors — broken contrast, a missing label — because errors deviate from consensus. But sameness doesn't deviate from consensus. Sameness is consensus. Ask three different frontier models whether a template-looking app is well designed, and you'll usually get three yeses.

The blind spot isn't a property of one model. It's a property of learned taste itself.

What actually escapes the loop

If averaged judgment can't see averaged output, you need judges that aren't built from averages. Two exist.

Math. Contrast is a ratio you can compute. Spacing consistency is a measurable pattern. Type scale relationships, touch target sizes, color usage — these have objective, checkable properties. A deterministic check doesn't consult taste at all; it measures, and measurements don't share a training set. A contrast failure is a contrast failure regardless of what any model finds aesthetically agreeable.

Human decision. The other escape is the one this blog keeps circling back to: deliberate choices no template would make. A committed palette. An opinionated detail. Unequal spacing that reflects how your content groups. The model can execute these decisions flawlessly — it just can't originate them, and it can't tell you they're missing, because their absence looks normal to it.

The practical workflow follows directly: let AI generate (it's brilliant at it), verify with measurement rather than vibes, and reserve the genuinely aesthetic calls for a human — you — armed with specific findings instead of a vague sense that something's off.

What this means for how Klyxx works

This blind spot is the design problem Klyxx was built around. A useful audit can't just be "AI looks at your screenshot and shares opinions" — that's the fox guarding the henhouse, the same averaged taste that generated the problem grading its own homework.

So a Klyxx audit is structured around the two escapes. Findings are anchored in a fixed rubric of sixteen dimensions with defined criteria — the same checks, applied the same way, every time — rather than whatever a model happens to notice on a given day. Where a property is measurable, it gets measured, not eyeballed. And the output isn't a verdict but a severity-ranked list of specific findings with fixes, because the final aesthetic call belongs to the person shipping the product. The audit's job is to make that call an informed one.

Get a judgment that isn't grading its own homework. Run your app through a structured audit — sixteen dimensions, severity-ranked findings, paste-ready fix prompts. Try Klyxx free.

The takeaway

AI builders are the best thing that ever happened to shipping speed, and the worst thing that ever happened to visual differentiation — for the same reason. The distribution that lets a model generate a competent interface in seconds is the distribution that makes every generated interface converge, and the model can't see the convergence because it has nothing else to see with.

That's not a reason to stop vibe-coding. It's a reason to stop asking the generator for a grade.

Get a free UX audit for your app or site

Try Klyxx free