PAPER A — FRAMEWORK · FIELD-TESTED · v02

The Legibility Audit: Testing AI artifacts with AI

Don't ask AI if the prototype is good. Ask what it thinks the prototype means.

WRITTEN AND DEVELOPED BY JEN SEPSO · MARCH 2026 · © 2026 JENNIFER SEPSO

POLISH IS NOT PROOF

Most conversations about AI and design begin with production: what can it make, how quickly, and how close can it get to something that resembles a finished product? The appeal is obvious. Screens that once took days can appear in minutes, complete with tables, filters, panels, empty states, and the posture of enterprise software. The artifact looks coherent. Sometimes it is even correct.

Those are not the same achievement.

While designing a planning workflow in a highly technical domain, I repeatedly encountered prototypes that looked mature before the underlying product logic was mature. Labels sounded credible. Components were aligned. Familiar patterns implied that somebody had thought everything through.

One early AI-generated prototype imposed a linear sequence on a workflow that needed to loop. It looked reasonable because linear progress is a familiar interface convention. Conceptually, it was wrong. The artifact had made the wrong mental model feel polished.

That changed the design question. Instead of asking whether AI could generate a screen, I began asking whether the resulting artifact could communicate the intended system without the designer standing beside it to explain.

THE WORKING QUESTION

A Legibility Audit is a small, structured way to test what a prototype appears to communicate. Instead of asking AI to critique the design, you ask it to interpret the artifact: give it the prototype, limit the context, and ask what it understands — what the product is for, what workflow it supports, what the system seems to be doing in the background, what persists, and what still feels unclear. Then you compare that interpretation against what you intended.

The AI is not the judge. It is the interpreter.

That distinction is the whole method. A critique asks: is this interface usable? A Legibility Audit asks: what do you think this interface is communicating? The judgment still belongs to the designer — whether the interpretation is accurate, whether the gap matters, and what should change because of it.

THE BOUNDARIES

The audit is not a heuristic review, a synthetic usability test, or a substitute for user research, domain expertise, accessibility evaluation, or stakeholder review. It has not met your users, your regulator, or the person who always finds the impossible edge case. Those methods answer different questions. The audit sits earlier. It is a clarity check before the larger review moments. It asks whether the artifact is already carrying the intended product logic or whether that logic exists only in the designer’s narration. This matters because AI-assisted production can make the interface look complete before the product has earned that appearance.

THE SETUP — SAME PROTOTYPE, DIFFERENT CONTEXT

I tested the same prototype in three clean AI sessions. The prototype stayed the same; the context changed. The separation mattered — if the model already knew what I meant because I'd corrected it earlier, the audit would be useless.

RUN 1
Prototype only
Screenshots. No explanation, no product description, no setup.
RUN 2
+ short description
One sentence: product category and user goal.
RUN 3
+ source workflow
The underlying flow representing the intended structure.
FIG 01 — THREE CONDITIONS, ONE PROMPT. Same questions every run — change the questions and you're no longer comparing interpretations, you're comparing prompts.

The prompt was deliberately plain: reconstruct the objective, workflow stages, objects, background processes, assumptions, and ambiguities. What I didn't ask: "how would you improve this?" That would have turned the exercise into critique. I wanted the model to stay in interpretation mode long enough to see the artifact through another lens.

WHAT CAME BACK

Across the three runs, the model reconstructed the core objective and major workflow stages with more alignment than I expected. The full-workflow run performed best — not surprising. The more interesting part: the prototype-only run was already fairly strong. Added context improved terminology and confidence but did not rewrite the interpretation. The labels, hierarchy, screen relationships, and visible parameters were carrying more semantic weight than I had realized.

That did not prove the design was good. It only suggested an outside interpreter could reconstruct the basic system logic from the artifact alone. A modest finding, but a useful one. The better part was where the interpretation slipped: the model inferred a more linear workflow than I intended, and raised questions about scale, persistence, and invisible background work. Those weren't embarrassing misses. They were the useful part. The gaps became the revision list.

THE GAP IS THE FINDING

A Legibility Audit is not a correctness machine. It is a mismatch finder. The question is not "was the AI right?" — it's "what in the artifact made that interpretation reasonable?" Maybe the navigation implies a false sequence. Maybe the labels are too generic. Maybe the object model is unclear, the system status invisible, or a familiar UI pattern is pulling the design toward the wrong mental model. In my case, the most useful finding: the prototype was quietly encoding a stronger sequence than I meant to imply. Exactly the kind of thing to catch before a stakeholder review, not during one.

I translate the audit into a simple revision table — more working note than research artifact. The output is not "AI feedback." The output is a decision: revise the structure, clarify the object model, change the labels, expose system status, or deliberately accept the ambiguity because it isn't important yet.

WHAT TO LOOK FOR
Purpose
Can it explain what the product helps the user do? If not, the core value proposition isn't communicating.
Stages, loops, decision points
This is where false linearity shows up.
Invisible system work
Background processing, validation, generation, saving. If it's invisible, users won't know what they're waiting for or what changed.
Persistence
Does it understand what carries across screens — a case, plan, scenario, task, decision? If not, the object model is unclear.
Confident wrong inferences
Often more useful than the things it gets right.
The questions it asks
Surface polish, or actual product logic? The best audit output isn't the one with the fewest questions — sometimes it's the one that asks the right ones.
WHEN TO RUN ONE

Run it when the prototype represents system behavior, not just static content: the product is complex or technical, the workflow includes loops or invisible background processes, AI helped generate the artifact, the design depends on labels and object relationships being understood — or you suspect the prototype only makes sense when someone narrates it. Don't use it for visual polish, preference testing, or final validation. It will not tell you whether users love the product. It is a pre-review clarity check: before I put this in front of other people, what does this artifact appear to say?

THE TEMPLATE
01 Write down what the prototype is supposed to communicate
02 Open a clean AI session — no prior project context
03 Upload the prototype or screenshots
04 Ask it to explain: objective, workflow, objects, background processes, assumptions, ambiguities
05 Repeat with one additional layer of context if needed
06 Compare the interpretation against your intended model
07 Turn the gaps into revisions
08 Save the evidence — especially the mismatches

The most important rule: do not correct the model mid-run. Let it misunderstand. The misunderstanding is the data.

THE TAKEAWAY

AI makes versions cheap, not decisions. The bottleneck moves from producing a first version to deciding which version deserves to continue.

A prototype is an argument about how a system works, and familiar patterns sometimes argue for workflows the designer never intended. AI is useful not because it is always right, but because its misunderstandings expose the artifact’s signals. If an outside interpreter cannot reconstruct the basic logic, the design may not be wrong. But it is not done communicating.

Written and developed by Jen Sepso · March 2026 © 2026 Jennifer Sepso

← INDEX NEXT: PAPER B →