Original research · 20 September 2026

What automated accessibility testing catches, and what it misses

Automated accessibility testing settles part of WCAG and leaves the rest to a person. On 20 September 2026 we ran axe-core over the home pages of 410 of the world's most-visited sites: 74% failed at least one check a machine can decide, and 81% also returned checks it could not decide alone.

Published · ShipperAG team · Data collected 20 September 2026, UTC

What we found

We scanned the home page of each of the top 800 domains in the Tranco ranking. 410 returned a page we could measure; 358 were not reachable as a normal web page, because they are content-delivery or API hosts rather than websites, and 32 were skipped because their robots.txt asks crawlers to stay out. Every number below describes those 410 pages.

  • 74% of pages (305 of 410) failed at least one WCAG 2.0, 2.1 or 2.2 A or AA check that axe-core can decide on its own.
  • 26% (105 pages) failed none of those checks. That is not the same as being accessible, as the next section explains.
  • The median failing page broke 2 rules across 3 elements; the worst broke 11 rules across 319 elements.
  • Of the rule failures we counted, 216 were rated critical and 572 serious by axe-core. There were no moderate or minor failures, because the A and AA rule sets we ran do not contain any.
  • 81% of pages (332 of 410) also returned at least one result axe-core flagged for a human to review rather than deciding itself.

These are the most visited and best-resourced sites on the web. They are not a fair picture of the average site, and that is the point: if 74% of them fail checks a machine can make in seconds, the checks are not the hard part.

On 2 October we ran the identical scan on 206 apps built with an AI app builder. The headline rate was similar, but the failures were very different: read the AI-built apps study.

What a machine decided

35 distinct rules failed somewhere in the sample. These twelve account for most of it.

Most common failures across 410 home pages, axe-core 4.10.3, WCAG A and AA rules only
RuleWhat it meansPagesShare
color-contrastText and its background are too close in colour to read comfortably14335%
link-nameA link has no text a screen reader can announce9824%
target-sizeA control is smaller than the minimum touch target size7318%
image-altAn image has no alternative text6315%
html-has-langThe page does not say what language it is written in, so a screen reader guesses the accent6215%
button-nameA button has no accessible name328%
aria-required-childrenAn ARIA role is missing the child roles it requires256%
aria-allowed-attrAn ARIA attribute is not allowed on that element256%
frame-titleAn iframe has no title246%
listA list contains elements that are not list items246%
nested-interactiveA control contains another control, which confuses assistive technology225%
link-in-text-blockA link is distinguished from surrounding text by colour alone184%

The shape is familiar to anyone who has run these tools: color-contrast leads at 35% of pages, link-name follows at 24%. 15% of pages did not declare a language on the <html> element, which takes one attribute to fix and changes how a screen reader pronounces every word on the page.

Every one of these is a rule a machine can settle without an opinion. They are also the cheapest possible class of defect: each one has a single correct answer, and none of them requires anybody to think about your product.

What a machine could not decide

axe-core does not only pass or fail. It also returns results it cannot settle, where the answer depends on meaning rather than markup. On 81% of the pages we scanned it did exactly that, on a median of 1 check per page and as many as 7.

That is the honest boundary of automation, and it is where the interesting failures live. A tool can tell you an image has alternative text. It cannot tell you the text describes the image. It can tell you a heading exists. It cannot tell you the heading describes what follows. It can tell you every control is reachable by keyboard. It cannot tell you the order makes sense to somebody buying something. Deque, who build axe-core, publish their own coverage analysis; the W3C's guidance on choosing evaluation tools states plainly that tools cannot determine accessibility on their own.

So a clean automated result is one sentence, not a verdict: no machine-decidable rule failed on this page as it loaded.

What this means before a release

Most teams treat an accessibility tool as a gate that is either green or red. The data says a release needs four answers, not two:

  • Verified. A check ran and passed, and you can say which check and when.
  • Issue found. A check ran and failed, reproducibly, with the element named.
  • Needs input. The tool reached its limit and a person has to judge: is this alternative text right, does this order make sense.
  • Not covered. Nothing checked it at all, which is the state most reports quietly omit.

The 81% of pages with results needing review are all sitting in the third state whether anyone writes it down or not. A report that shows only pass and fail turns that third state into a silent pass, which is how an accessible-looking release reaches a user who cannot use it. If you want a checklist to run this yourself, our website QA checklist marks which checks a tool can decide and which need a person, and AI accessibility testing covers the WCAG 2.2 side in more depth.

Method

  • Sample. The top 800 domains of the Tranco daily top-1M list generated on 19 September 2026, downloaded and scanned on 20 September 2026 UTC. Tranco is a research-oriented ranking built to be reproducible and hard to manipulate.
  • Pages. One page per domain, the home page at https://<domain>/, following redirects. No other page was requested, and nothing was submitted. The run was restarted once after a site hung the scanner, so a minority of domains were requested a second time; where that happened we kept a single result per domain, preferring the attempt that returned a measurable page.
  • Browser. Chromium through Playwright at 1440 by 900, reduced motion, waiting for the document and then 2.5 seconds for late-loading content.
  • Rules. axe-core 4.10.3, restricted to the wcag2a, wcag2aa, wcag21a, wcag21aa and wcag22aa tags. Best-practice rules were excluded deliberately: they are good advice, not the standard.
  • Politeness. robots.txt was read first and any site disallowing all crawlers was skipped. One request per site, spaced out, no logins, no forms, no repeat visits.
  • Counting. A "failing page" is a page with at least one violated rule. "Pages" in the table counts pages with at least one instance of that rule, not instances.
  • Privacy. We publish aggregates only. No per-site result is published here or anywhere else.

What this study is not

It is not a WCAG audit of anybody. A conformance claim needs a person, several assistive technologies and more than one page, and none of that happened here. It is not a ranking, and we do not name which site did what. It measured one page per site on one day, in one browser, at one window size, in the state the page happened to load in: consent banners, region redirects and lazily loaded content all change what a scanner sees. And it is a snapshot of the largest sites on the web, so it says nothing about the average site.

The comparison people will reach for is the WebAIM Million, which found detected WCAG 2 failures on 95.9% of a million home pages in February 2026. That uses a different engine, a wider rule set and a sample a thousand times larger. The numbers are not comparable. They point the same way.

Use the data

The rule-level results are published as a CSV under CC BY 4.0: automated-accessibility-study-2026.csv. Credit ShipperAG and link to this page. If you repeat the scan with the method above you should get numbers close to ours, and we would like to hear if you do not. The scanner we use for this method is published under the MIT licence: accessibility-study-scanner.mjs. Related reading: AI testing statistics for sourced figures from other people's research, European Accessibility Act testing for what the law now requires, and the glossary for any term above.

To cite it, copy this:

ShipperAG (2026). What automated accessibility testing catches, and what it misses. An axe-core scan of 410 high-traffic home pages, 20 September 2026. https://shipperag.com/automated-accessibility-testing-study/

Sources, all opened and checked 20 September 2026: Tranco, a research-oriented top sites ranking; Deque, axe-core and The Automated Accessibility Coverage Report; W3C WAI, Selecting Web Accessibility Evaluation Tools and WCAG 2.2; WebAIM, The WebAIM Million 2026.

// FAQ

Automated accessibility testing, answered

What can automated accessibility testing actually detect?

Automated tools decide the rules that can be judged from the page's code: colour contrast, missing alternative text, a link or button with no accessible name, a missing language attribute, controls that are too small to tap. In this scan axe-core reached a verdict on a median of 26 checks per page. It cannot decide whether alternative text is accurate, whether a heading describes the section below it, or whether a keyboard path through a checkout makes sense. Those need a person.

What share of sites fail automated accessibility checks?

In this scan, 74% of the 410 home pages failed at least one WCAG 2.0, 2.1 or 2.2 A or AA check that axe-core can decide on its own, and 26% failed none. The most common failure was color-contrast, on 35% of pages. These are the most visited, best-resourced sites on the web, which is what makes the number worth knowing.

Why is this number lower than the WebAIM Million?

Because it measures something narrower. WebAIM's analysis of a million home pages in February 2026 found detected WCAG 2 failures on 95.9% of them, using its own WAVE engine across a wider rule set and the whole top million. This scan used axe-core with the WCAG A and AA rule sets only, no best-practice rules, on a much smaller and more heavily resourced sample. Different tool, different rules, different sample: the two numbers are not comparable, and both point the same way.

Does passing an automated check mean a page is accessible?

No. A clean automated result means no machine-decidable rule failed on the page as it loaded. It says nothing about the checks a tool cannot decide, and in this scan 81% of pages also returned at least one result that axe-core flagged for human review. The W3C is explicit that evaluation tools cannot determine accessibility on their own.

Can we reuse this data?

Yes. The rule-level results are published as a CSV under CC BY 4.0: credit ShipperAG and link to this page. The method below is written out in full so anyone can repeat the scan, and we publish aggregates only, never the results for a single site.

// free 45-day trial · waitlist open

Four answers, not two

ShipperAG checks each release against what it is supposed to do and reports what was verified, what failed, what needs a person and what was not covered. It is in a private pilot for tech and product teams.

We keep what you enter here and the page you joined from, only to contact you about the ShipperAG pilot. No spam.