Skip to main content

Method

How Personaudit tests, and what it does not claim

Last reviewed 10 October 2026

On this page

This page states how a Personaudit result is produced, what our own experiments found, and where they found nothing. The experiments are public, with their pre-registered thresholds, so you can read the raw write-ups.

The scan engine

Accessibility findings come from axe-core 4.13.0, run through Playwright. The scan makes no model calls, so the same page state produces the same findings every time.

Each state the crawl reaches is checked against these axe tags: wcag2a, wcag2aa, wcag21a, wcag21aa, wcag22a, wcag22aa, and best-practice.

A defect is identified by its axe rule plus a normalized element selector. The normalization collapses ids that a framework generates on every render, so one broken component reads as one defect across every state, not one per page.

Scanning behind a login runs in the CLI on your machine, using a session you save yourself. The hosted service does not scan behind a login.

What we measured

Three experiments shaped what Personaudit claims. Each wrote its metric and kill criterion down before the data existed. One of them killed a claim we had planned to make.

Do deep states hold violations a page scan misses?

Source: experiments/net-new-violations. Runs on 15 and 16 July 2026, against vendor demo apps and a local Metabase fixture, not production sites.

  • On Sauce Demo, a scan of the entry page found 0 blocking violations. Behind the login there were 3 critical ones.
  • On the-internet, deep states added 5 net-new findings to 44 already on the entry page. That run missed the threshold set in advance.
  • In the four-target run, the entry-page verdict was wrong on 1 of 4 targets. On applitools-demo, deep states hid nothing (0 net-new findings).
  • On the signed-in Metabase fixture, 66 of 85 affected elements (78%) appeared only in states deeper than the entry page.

The write-up says its own metric changed between runs, that there are only a few targets, and that one OrangeHRM comparison is a post-hoc observation, not a pre-registered result. It also says this measures reach, not WCAG coverage, and that a scripted crawler with a saved session would have reached the same states.

Do personas find accessibility defects a crawler misses?

Source: experiments/personas-vs-crawler. One target, a Metabase fixture, run on 16 July 2026. No. The pre-registered metric came out at 13.7%, below the 15% line that marks the claim as dead.

  • The crawler reached 40 states and the personas reached 9.
  • The crawler found 264 defects and the personas found 92.
  • The two shared 50. The crawler alone found 214, the personas alone found 42.
  • The crawler used no model calls. The personas used about 90.

The defect key in that run was flawed: framework-generated selectors made the same element look like two defects. The write-up records the flaw instead of rescoring, and a sensitivity check that normalized the ids moved the figure to 11.0%, which does not change the verdict. The only thing personas reached that the crawler structurally could not was content inside opened menus: 3 defects out of a union of 109.

So Personaudit does not claim that personas find accessibility defects. The deterministic scan does that.

Can you trust a persona's task-success verdict?

Source: experiments/task-success-validity. Two targets, 10 goals each (5 reachable, 5 impossible), every goal run once, on 16 July 2026.

  • Metabase: 9 of 10 verdicts matched the known answer. None of the 5 impossible goals was reported as achieved.
  • Sauce Demo: 9 of 10 matched, again with none of the 5 impossible goals reported as achieved.
  • Each run had one false "blocked" on a goal that was reachable: a dashboard the persona could not find, and a native select control it could not operate within its step budget.

That is the error profile: when a persona reports "blocked", the site may be blocking the goal or the persona may have failed to find a way. The write-up calls the result enough to build on, not a launch claim, because the sample is small and single-run.

What personas do

Personas browse toward a goal and report two things: whether the goal was achieved, and an opinion about the experience. They do not issue an accessibility verdict and they never simulate a disabled user. A test in the repository fails the build if a persona profile claims a disability or is asked to judge WCAG conformance.

Traversal personas are told they are an automated harness, not a person. For testing with disabled people, see the accessibility statement.

What automation cannot see

Automated checks, ours included, catch only part of the failures WCAG describes. A clean scan means axe found no violations in the states it reached. It does not mean the site works for everyone. You still need manual keyboard testing and review with a screen reader.

For how we check this site against the same engine, see the accessibility statement. For how we handle your data, see the security page. For what a finished report looks like, see the sample report.