Quality lead reviews AI quality standards on a dashboard beside printed design guidelines

Stripe’s Katie Dill has a name for what happens when AI makes building fast and easy: “zombie UI.” Software that looks finished, passes review, and has no point of view behind it. Her talk at the Lenny and Friends Summit has been watched nearly 250,000 times, mostly by product and design people. It should be required viewing for quality and operations leaders, because the same failure mode is already showing up in your processes. AI-generated deviation reports, CAPA write-ups, and audit summaries that tick every box and tell you nothing.

Speed without judgment produces output that satisfies the form and fails the customer. Below, how to set AI quality standards that encode your team’s actual point of view, where to put the human checks, and what it changes in your numbers.

Diagram: AI Quality Standards: Scale Craft Without Generic Output
Process diagram — AI Quality Standards: Scale Craft Without Generic Output

Everything Your AI Produces Passes Review and Nothing Stands Out

Dill’s argument is that good work still depends on three things AI does not supply on its own: a clear point of view, an understanding of the people using the output, and care in the details. Strip those out and you get volume. Work instructions that read like a manual for a different factory. Audit responses that answer the question without answering the finding. Supplier emails that are polite, complete, and forgettable.

The cost of producing a document used to be the control mechanism. Writing took hours, so someone had to decide what mattered and defend it. That friction is gone.

Nobody fails an internal review for being generic. That is why AI output quality degrades quietly, and why your own standards have to carry the weight the effort used to.

Rows of nearly identical app screens on a wall, each marked as meeting AI quality standards

What ‘Zombie UI’ Actually Describes, and Its Equivalent on the Factory Floor

A zombie interface works. Buttons click, forms submit, nothing errors out. It was built fast because the tool made building cheap, and nobody decided what it was for. The equivalent in operations is an 8D report that recites the template, names a containment action, and never diagnoses a root cause. Or an SOP no operator on the line can actually follow.

The three symptoms: no point of view, no user understanding, no care in the details

No point of view shows up as output that hedges. The deviation is “under investigation,” the risk is “moderate,” the recommendation is “continue monitoring.” Nothing in the document commits to a position someone could be wrong about.

No user understanding shows up as a dashboard built for a reviewer rather than a decision-maker. Twelve charts, no threshold, no owner. The diagnostic question is simple: would a competent human reading this know what to do next? If not, you produced a document, not a decision.

Why ‘it passed the checklist’ is the weakest possible quality signal

Checklists verify presence, not substance. An AI model is extremely good at producing presence. Every field populated, every section headed correctly, every reference cited.

So “it passed review” now tells you almost nothing about AI output quality. Your acceptance criteria were written when generating text was expensive. They assumed effort. Drop that assumption and the checklist becomes a formatting test.

The Three Inputs AI Cannot Supply for You

A point of view, an understanding of who consumes the output, and care in the details. Dill frames these as the things that survive when building gets cheap. For a quality or operations function, they are not attitudes. They are documents, and if nobody has written them, the model has nothing to work from.

Turning tacit quality judgement into written standards

Your best quality engineer knows which deviations warrant a full investigation and which ones close out in a paragraph. That judgement lives in their head, refined over fifteen years of audits. An AI model cannot infer it from your template library.

So write it down. What a sufficient root cause looks like in your plant, which three details an operator needs before a line restart, what tone you take with a supplier who missed spec twice. These become the prompt layer, the review criteria, and the training material at the same time. Teams that have never articulated their standards cannot scale them, they can only scale their formatting.

Who in your organisation owns the point of view

Someone has to hold the pen. In practice this fails when it gets handed to IT or to whoever ran the pilot, because neither of them owns the outcome the document produces.

Give it to the person who signs off on quality. They define what good looks like, approve the standard, and review samples monthly. Not a committee. One accountable name.

Diagram showing three human inputs feeding AI quality standards: viewpoint, user insight, and craft

Building Your Standards Into the AI Tools, Not the Review Step

Dill’s actual recommendation is the part most people skip: build the standards into the tools. Not into a checklist someone runs afterward. The difference decides whether AI frees up bandwidth or quietly consumes more of it.

From review gate to encoded guardrail

A human reviewer at the end is the default pattern because it requires no setup. It also fails at volume. Three reviewers cannot absorb thirty times the document output, and the ones who can judge quality are the people you were trying to free up.

Encoding means the standard sits upstream. Your written criteria become the system prompt. Your approved SOPs, specs, and prior investigations become the retrieval source, so the model cites your plant and not the internet. Structured templates force the fields that matter to be filled with something specific.

Writing evals for quality documents: what to score and how

Dianne Penn of Anthropic put it well in a talk titled “Evals are the new PRDs.” The specification of good becomes the deliverable. Write the test before you write the prompt.

Score narrowly and name the criteria. Does the root cause statement identify a mechanism, not a category? Is every corrective action tied to a named owner and a verification method? Does it reference the actual revision of the spec in force?

Run those checks automatically. Anything failing never reaches a human.

Where AI Scales Quality and Where It Quietly Erodes It

The split is cleaner than most vendors admit. AI handles work where the correct answer is already defined somewhere and can be checked against a reference. It degrades work where someone has to decide what the correct answer is.

The explicit-standard test

Ask one question before automating any process: could two trained people independently produce the same output and agree it was right? If yes, the standard is explicit and AI will scale it well. Defect classification against a coded taxonomy. Deviation triage using documented severity criteria. Rewriting an approved SOP into language the line actually uses. First-pass audit evidence assembly, where the task is finding and labelling records, not arguing about them.

These share a trait. The output is verifiable against something written. You can test it, measure the error rate, and tighten the prompt. That is how you get quality at scale rather than volume at scale.

Processes to leave in human hands for now

Root cause analysis fails the test. So does supplier escalation, regulatory interpretation, and any decision that becomes precedent for the next hundred cases. The standard here is not written down because it depends on context that changes case by case.

Dill’s framing is that quality depends on “care in the details.” Where the details carry consequence, keep a named person accountable for the judgement. Let AI assemble the evidence, draft the timeline, surface the contradictions. Then let someone decide.

Split-panel diagram contrasting where AI strengthens AI quality standards versus where oversight erodes

Ready to find AI opportunities in your business?
Book a Free AI Opportunity Audit. It is a 30-minute call where we map the highest-value automations in your operation.

A 90-Day Plan to Encode Quality Standards Before You Scale Volume

Pick one document type with real volume. Deviation reports, incoming inspection notes, supplier corrective action requests. One. Trying to encode three at once means encoding none of them properly.

The artefacts you should have at the end of each 30-day block

Days 1 to 30 produce a written standard, built from twenty examples you have graded. Ten that your best engineer would sign without edits, ten that would get sent back. The standard is the articulated difference between the two piles. If you cannot write that difference down, you have found the real problem.

Days 31 to 60 produce three things: a prompt, a template, and an eval set of roughly thirty held-out cases with known correct answers. Run the model against the evals and record a pass rate. Days 61 to 90 produce a routing rule. Output that clears the eval goes to a human for sign-off. Output that fails gets regenerated or escalated, never reviewed.

What to measure so the gain survives contact with finance

Three numbers. Reviewer hours per hundred documents before and after. Rework rate, meaning documents returned for revision. Audit findings tied to documentation quality, measured across a full cycle rather than a quarter.

Pass rate becomes the number you manage week to week. It moves when the standard improves or the process drifts, and it tells you which. Report hours returned, not accuracy percentages. Finance buys capacity.

Craft Becomes the Advantage When Everyone Can Build Fast

When producing a document costs almost nothing, the document stops being the output. What matters is the thinking in front of it and the inspection behind it. Everyone on your competitive set now has access to the same models at the same price. The gap between two plants running identical tooling comes down to whether anyone bothered to define what good looks like before switching the volume on.

Dill’s framing is that quality depends on “a clear point of view, an understanding of users, and care in the details.” Those three things compound. An organisation that writes down its judgement this quarter has a reusable asset next quarter, and a sharper one the quarter after. An organisation that scales volume without that groundwork is just producing mediocrity at a faster rate, with more of it to unpick later.

The quality leader’s new job description

Your job used to be reviewing work. It is now defining the standard that work is measured against, in language precise enough that a model can apply it without you in the loop. That is editorial work, not administrative work, and it belongs to your most experienced people.

Take one question into your next leadership meeting. Could we hand our definition of quality to a machine, and would we recognise the result as ours? If the answer is no, you have your priority, and it is not buying another tool.

Source: youtube.com

Leave a Reply