Who sets the quality bar?
On what’s missing in most AI product teams, and a template to start fixing it.

Designer Fund just published their AI in Design 2026 report. Over 900 designers surveyed across over 60 countries. The most rigorous study of AI and design practice we have.
There is something in the data the report doesn't call out directly: most AI products don't have a quality bar. The gap between the demo and the lived user experience is almost always a design gap, and the reason it exists is usually that nobody set a standard before building started.
What the data says
The top challenges designers cite when using AI are unreliable and inconsistent output quality, lack of control over output, and lack of product and brand context.
The AI doesn't know what this product is trying to do, what good looks like for this user in this moment, what the intent is behind the thing being built. Those things improve when someone specifies them before building starts, not by trying harder with better prompting.
The report also finds that collaboration has decreased. In 2025, 5% of designers said AI had reduced collaboration with teammates. In 2026, that number is 20%. Designers are working more alone, more in terminals and prompts, less in the conversations where intent gets surfaced and shared.
Put those findings together and you get a picture of teams producing more output, faster, with less shared understanding of what good looks like. That shows up in the product. And in the experience of the person trying to use it.
Why it’s hard
Speed breaks the traditional model. The quality bar used to be applied through critique, review, and sign-off. That process assumed some time between generation and shipping. When prototypes move from prompt to production in hours, the bar has to be embedded before anyone starts, not applied after.
Visual quality is much better, which is part of what makes this hard to see. AI makes everything look polished. Things can pass a visual bar while failing completely on whether they serve the intent, communicate uncertainty honestly, or behave correctly in the moments that matter. The old quality bar was partly visual. The new one has to go deeper, and the failure modes are invisible until someone tries to use the product.
Ownership is genuinely unclear. When a PM prompts a feature into existence, a designer reviews it, and an engineer ships it, who held the bar? Each of them applied some judgment. Nobody called it on the whole experience.
The critique problem
The traditional mechanism for maintaining the quality bar was critique. A designer brings work to a room. The work represents days or weeks of effort. The conversation asks: is this the right direction, is the craft good, does this serve the user. The bar gets held through that exchange.
That model has changed shape, and most teams haven’t noticed yet.
Volume is the first problem. AI can generate fifty directions in an hour. You cannot examine fifty directions the way you’d examine five. Something shifts when the process becomes pattern-matching across many artifacts rather than examining each one carefully.
The question itself has changed. Old critique asked: is this right? New critique has a second layer nobody has fully worked out yet. What did the prompt ask for versus what did we get? What did the human catch versus what slipped through? You’re evaluating the human-AI collaboration as much as the artifact. That’s a different skill, and most teams are still using the old one.
Then there’s the speed mismatch. Work ships faster than critique can happen. The review is often retroactive, after the thing is already in production. Critique can no longer function purely as a quality gate. It becomes a calibration tool. Every “this is wrong” should update the encoded criteria so the same mistake doesn’t repeat in a different prompt next week. Otherwise the insight evaporates, and nothing changes.
Who sets the quality bar?
Right now, in most AI products, nobody does.
The report describes something important: designers are beginning to pre-program design system components and brand guidelines into coding tools so that anyone generating interfaces starts from a shared quality baseline. That is intent specification work. That is the quality bar being established before anyone picks up the tools. It’s incomplete. It’s imperfect. It’s directionally right.
You write it down before anyone starts building. The quality bar has to exist in a format that can be applied, not just understood in the head of the most experienced person on the team.
You make it specific enough to arbitrate a real decision. “Good design is clear” doesn’t help. “In this product, clarity means this over that, because our users are in this situation” is encodable. The specificity is the work, and it’s harder than it sounds because most teams have never had to write down what they actually believe about quality.
You own the not-doing list. The quality bar includes what the product refuses to do, what patterns it won’t adopt, what experiences are out of scope. When AI can produce anything, the bar is partly defined by what you’re saying no to.
Design has the methods to do this. The vocabulary of intent, behavior, and user need. The discipline of asking what something should do before asking how it should look. The practice of specifying what good looks like for a real person in a real situation, before anyone starts building.
But only if design is in the room when those questions get asked.
When nobody answers them, you get AI that technically works and consistently disappoints. The products are everywhere. You’ve used them. The question is whether your team is building one.
👉 Download quality-bar.md - a 10-section template for encoding your quality criteria before you build. MIT licensed, free to use and adapt.

