Hide the model names, give reviewers the same practical task and ask what happened. Record whether they could use the design before asking which version they preferred. Reveal the generator only after those judgments are saved.
This is a proposed comparison method for agent-generated interfaces, reviewed September 12, 2026. No blind study was conducted for the article. The aim is a better local buying decision, not a new model leaderboard.
- 01Hide identity until after judgment.Keep an independent mapping from artifact labels to model and run.
- 02Separate function from taste.An attractive interface can fail the task or exclude users.
- 03Compare workflows fairly.Record the brief, budget and revisions that produced each candidate.
01 — Practical decisionWhy hide the model name at all?
The Chatbot Arena paper describes anonymous, randomized pairwise comparisons with identities revealed after voting. That is evidence of a preference-evaluation design for conversational systems. It does not establish that the same procedure has been validated for choosing business interfaces.
Our proposed adaptation is narrower: reduce the information about the generator that reviewers receive before judging an artifact. A famous model, a costly tier or a favored vendor can otherwise become part of the argument before anyone tries the interface. Hiding the label does not remove every bias, and visual habits may still suggest a source.
The comparison unit should be a delivered design with a defined task. Do not ask reviewers to choose the best AI model in general from two pages. Ask which candidate better supports the task you actually need, and what evidence supports that choice.
02 — Practical decisionFix the brief and identify the actual treatment
Give every candidate the same audience, content, brand constraints and user journey. Supply the same permitted assets. Specify what must work, what may be simulated and what must remain editable. The design-file versus screenshot guide helps define those inputs.
Choose whether you are comparing equal spending, equal elapsed time or each workflow’s normal production process. None is automatically fair for every question. Equal prompts with radically different revision budgets do not isolate a model difference. Normal production workflows may deliberately include different tools, but that needs to be the declared comparison.
Preserve the generated source and every revision selected for review. Avoid presenting one candidate’s first draft against another candidate’s carefully repaired final version without saying so. The table below separates three judgments that should remain visible throughout the decision.
| Judgment | Question | Evidence to retain |
|---|---|---|
| Task success | Can the person complete the intended action? | Observed result, blockers and assistance. |
| Accessibility | Can relevant users operate and understand it? | Human checks plus scoped tool findings. |
| Visual preference | Which presentation better serves the brief? | Recorded preference with concrete reasons. |
| Delivery cost | What did accepted output require? | Attempts, repair time, charge and editable files. |
03 — Practical decisionPrepare anonymous artifacts that remain usable
Assign neutral labels to artifact versions and store the model/run mapping away from reviewers. Remove visible model branding and obvious generator labels where doing so is lawful and does not alter the product being tested. Keep required credits and licensing information intact.
Vary presentation order across reviewers and record it. A reviewer seeing one candidate first may use it to learn the task, making the second easier. For interactive work, use equivalent starting state, data and browser conditions. Do not share login state or a completed form between candidates.
If a reviewer recognizes a likely model from its style, let them record that suspicion without confirming it. Complete blinding may be impossible. The useful record states what was hidden, what could still identify the generator and whether the judgment was committed before reveal.
04 — Practical decisionAsk people to complete a task before scoring appearance
Give a concrete task: find a service price, compare two options or recover from a failed form submission. Record the observed outcome and where the reviewer needed help. These are proposed task examples, not results. Avoid invented stopwatch figures or a single overall score with unexplained weights.
Check keyboard use, labels, focus visibility, readable text and relevant responsive states. W3C’s accessibility overview states that no tool alone determines whether a site meets accessibility standards; knowledgeable human evaluation is needed. A blind preference vote is likewise not an accessibility certification.
After task checks, ask about hierarchy, clarity and visual preference. Require a short explanation tied to the artifact. A comment that the pricing distinction is difficult to find is more actionable than saying a candidate feels less premium. Keep severe functional failures visible even if the same design wins the taste vote.
05 — Practical decisionMake ties and disagreement useful evidence
Allow a tie and an insufficient-evidence response. Forcing a choice can create a winner that nobody meaningfully prefers. If reviewers disagree, inspect whether they performed different tasks, used different devices or valued different parts of the brief. Do not average away a failure affecting a critical audience.
An illustrative result could favor one candidate for navigation and another for editing existing content. That suggests a revision or a workflow choice, not necessarily a single winner. Keep the evidence by task so the final decision can reflect the business’s actual priority.
The editable-export reference covers what happens after selection. A design that reviewers like may still be costly to maintain if the team cannot edit its source. Treat that as a separate delivery criterion and disclose it before making the purchase decision.
06 — Practical decisionReveal the labels and compare the cost of accepted delivery
Freeze the reviewer record before revealing model names, pricing or generation time. Then bring in the operational evidence: attempts, cost, repair effort, editable files and any required dependencies. This sequence keeps those facts available for procurement without allowing them to substitute for direct design judgment.
A small internal comparison supports a decision about the tested brief and workflow. It does not establish statistical superiority across products or customers. If the difference is slight, prefer a reversible choice or gather more representative cases rather than turning a fragile preference into a public claim.
Use the interactive-demo guide to make the review artifact executable. In an AI transformation project, the deliverable should include both the chosen design and the evidence explaining why it fits the job.
Evidence and scope
- As-of date
- September 12, 2026: sources retrieved and reviewed. September 12 is the editorial allocation. Verified event dates are stated separately.
- Sources and method
- Original decision procedure informed by the Chatbot Arena paper and W3C accessibility evaluation guidance. Sources checked September 12, 2026.
- Limits
- No reviewers recruited, candidates generated or scores collected. Model anonymity reduces one influence; it does not prove absence of bias.
07 — Next stepChoose the design on evidence, then choose the workflow
Choose the design on evidence, then choose the workflow
Anonymous review can make the artifact the focus of attention. Keep task success, accessibility, preference and delivery cost visible as separate judgments. Reveal the generator after review, then choose the workflow that produces an acceptable result under your real constraints.