A shop is built with a defect planted in it, a population is sent through, and only afterwards is what happened compared against what was planted. Nothing upstream of the scorer ever sees the answer. This page is that test, its results, and the reason there is no headline percentage on it.
Each case is one published criterion, violated once, in one place. The population walks the shop, the detector clusters what happened to them, and the score is worked out afterwards by comparing the clusters against the page the defect was planted on. Detection means the struggle localised to the right place — the detector is never asked to name the cause, because it cannot and we would not believe it if it did.
There is always a clean arm: the same shop with nothing wrong with it. Every cluster it produces is the detector inventing a problem for somebody to chase, so a detection rate without it says nothing. A cluster on a defect arm only counts if it affects at least two more people than the clean shop raises in the same place.
Both columns below are the identical code with the identical settings, run twice with eight simulated people per arm. 7 of 14 cases landed somewhere different the second time. 4 of those are the detector reaching a different verdict; the other 3 are arms that ran short of usable sessions, which is the harness rather than the detector and is reported separately for that reason.
| Case | Criterion | Run 1 | Run 2 |
|---|---|---|---|
image-only-link not blind |
WCAG 2.2 SC 1.1.1 Non-text Content (Level A) | no effect | missed |
ambiguous-links |
WCAG 2.2 SC 2.4.4 Link Purpose (In Context) (Level A) | no effect | no effect |
silent-submit |
Nielsen heuristic #1 Visibility of System Status; WCAG 2.2 SC 4.1.3 Status Messages (Level AA) | found | no effect |
no-input-hint not blind |
WCAG 2.2 SC 3.3.2 Labels or Instructions (Level A) | found | found |
vague-error |
WCAG 2.2 SC 3.3.1 Error Identification and SC 3.3.3 Error Suggestion (Level A); Nielsen heuristic #9 | missed | found |
no-exit |
Nielsen heuristic #3 User Control and Freedom | no effect | no effect |
recall-code |
Nielsen heuristic #6 Recognition Rather than Recall | found | found |
redundant-entry |
WCAG 2.2 SC 3.3.7 Redundant Entry (Level A) | found | found |
disabled-no-reason |
Nielsen heuristic #1 Visibility of System Status; WCAG 2.2 SC 3.3.2 Labels or Instructions (Level A) | no effect | missed |
late-cost |
Nielsen heuristic #1 Visibility of System Status | no effect | no effect |
wrong-order-error |
WCAG 2.2 SC 3.3.1 Error Identification (Level A); Nielsen heuristic #9 | no effect | no effect |
offscreen-cta needs eyes |
WCAG 2.2 SC 1.4.10 Reflow (Level AA) | no effect | too small |
target-too-small needs eyes |
WCAG 2.2 SC 2.5.8 Target Size (Minimum) (Level AA) | no effect | too small |
inconsistent-naming |
WCAG 2.2 SC 3.2.4 Consistent Identification (Level AA); Nielsen heuristic #4 | no effect | too small |
no effect means the planted defect never made anybody measurably struggle — a fact about the fixture rather than a verdict on the detector. missed is the real failure: people demonstrably struggled and we could not say where.
no-input-hint, recall-code, redundant-entry.The instrument is not precise enough to rank two versions of the product against each other, which is what we most want it for. That is a sample-size problem and a threshold problem, not an interpretation problem, and no amount of reading the current numbers harder will fix it.
We publish this now, at this state, because a tool that measures other people's products should be readable about its own. If that is the only thing on this page you take seriously, it is the right one.