Reference · 21 September 2026

The 32 WCAG checks no tool makes for you

Automated accessibility testing covers less than half of WCAG, and not the hard half. axe-core 4.10.3 carries rules for 23 of the 55 success criteria at Level A and AA in WCAG 2.2. Here are the other 32, with what each one actually needs from a person.

Published · ShipperAG team

The count

axe-core 4.10.3 ships 69 rules in its WCAG A and AA tag sets. Those rules carry WCAG tags for 23 of the 55 success criteria that WCAG 2.2 defines at Level A and AA. The remaining 32 have no rule tagged to them at all.

That is not a criticism of axe-core, which is the best engine of its kind and is careful about what it claims. It is the shape of the problem. The criteria a machine can settle ask whether something exists: does this image have alternative text, does this control have a name, does this text meet a contrast ratio. The criteria below ask whether something makes sense, and meaning is not in the markup.

Pair this with the other half of the picture: in our scan of 410 high-traffic home pages, 74% failed at least one check a machine can decide, and 81% also returned checks it would not decide. Both numbers point the same way. The automatic part is cheap. The judgement is the work.

The 32, and what each one needs from a person

WCAG 2.2 Level A and AA success criteria with no rule tagged to them in axe-core 4.10.3
CriterionNameLevelWhat a person has to do
1.2.3Audio Description or Media Alternative (Prerecorded)AWatch the video. Decide whether the audio description or the text alternative actually conveys what is on screen.
1.2.4Captions (Live)AAAttend or sample the live stream and judge whether the captions keep up and are accurate.
1.2.5Audio Description (Prerecorded)AAWatch with the audio description on and check it covers what matters visually.
1.3.2Meaningful SequenceATurn off CSS, or read the page with a screen reader, and check the reading order still makes sense.
1.3.3Sensory CharacteristicsARead every instruction and look for ones that rely on shape, size, position, sound or colour alone: “click the round button on the right”.
1.4.5Images of TextAALook for text baked into images. A tool cannot tell a logo from a paragraph in a picture.
1.4.10ReflowAASet the window to 320 CSS pixels wide and check nothing needs two-dimensional scrolling and nothing is lost.
1.4.11Non-text ContrastAACheck that the parts of a control that carry meaning — its border, its icon, its focus ring — have 3:1 contrast. Which parts carry meaning is a judgement.
1.4.13Content on Hover or FocusAAHover and focus every tooltip and popover. Check it can be dismissed, hovered over, and stays until dismissed.
2.1.2No Keyboard TrapATab through the whole page, including every widget and modal, and check you can always get out again.
2.1.4Character Key ShortcutsALook for single-character shortcuts and check they can be turned off, remapped, or only fire on focus.
2.3.1Three Flashes or Below ThresholdAWatch any animation or video for flashing. Measuring it needs a person and, for anything borderline, a tool like PEAT.
2.4.3Focus OrderATab through the page and judge whether the order preserves meaning and operability. Only a person knows what the intended meaning was.
2.4.5Multiple WaysAACheck there is more than one way to reach each page: navigation, search, a sitemap.
2.4.6Headings and LabelsAARead the headings and labels and judge whether they describe what follows. A tool sees that a heading exists, not whether it is honest.
2.4.7Focus VisibleAATab through and look. Check the focus indicator is visible on every control, against every background it lands on.
2.4.11Focus Not Obscured (Minimum)AATab through and check nothing sticky — a header, a cookie bar, a chat bubble — hides the focused control.
2.5.1Pointer GesturesAFind every gesture that needs a path or multiple fingers and check there is a single-pointer alternative.
2.5.2Pointer CancellationAPress and hold on controls and check the action fires on release, and can be aborted by moving away.
2.5.4Motion ActuationAShake, tilt or move the device and check any motion-triggered action has a conventional alternative and can be switched off.
2.5.7Dragging MovementsAAFind every drag interaction and check there is a way to do the same thing with a single tap or click.
3.2.1On FocusATab into every control and check nothing changes context on focus alone.
3.2.2On InputAChange every input and check nothing submits, navigates or reorders the page without being asked.
3.2.3Consistent NavigationAACompare navigation across pages and check the order is consistent. That needs more than one page, which a single-page scan never sees.
3.2.4Consistent IdentificationAACompare icons and controls that do the same job across the site and check they are named the same way.
3.2.6Consistent HelpACheck that help — contact details, a chat, a help link — appears in the same relative place on every page that has it.
3.3.1Error IdentificationATrigger every error and check the message says which field is wrong, in text.
3.3.3Error SuggestionAATrigger every error and judge whether the suggestion is useful and correct.
3.3.4Error Prevention (Legal, Financial, Data)AAWalk through anything legal, financial or destructive and check it is reversible, checked, or confirmed.
3.3.7Redundant EntryAComplete a multi-step process and check it does not ask again for information you already gave.
3.3.8Accessible Authentication (Minimum)AATry to sign in without a cognitive function test, or check the alternative or the object-recognition exception applies. Password managers must be able to fill it.
4.1.3Status MessagesAATrigger status messages — “item added”, “3 results”, “saving” — and check a screen reader announces them without moving focus.

The pattern worth learning

Read the right-hand column and the same four verbs keep appearing: tab through, trigger, compare, judge. Three of those are mechanical. Only the last one needs a person.

  • Tab through covers focus order, keyboard traps, focus visibility and focus obscured. Somebody or something has to actually walk the page with a keyboard.
  • Trigger covers every error-handling criterion. You cannot check an error message that never fired, which is why error states are the most commonly missed part of an audit.
  • Compare covers consistent navigation, consistent identification and consistent help, and it needs more than one page. A single-page scan is structurally incapable of it.
  • Judge is the residue: is this heading honest, is this suggestion useful, does this order preserve meaning. That is the part no tool takes off you, and the part worth your attention.

An agent can do the first three and hand you the fourth with the evidence attached. That is the whole argument for agentic testing, and it is why our reports carry four states rather than two: verified, issue found, needs input, not covered. A criterion nobody checked is not covered, and saying so is more useful than a green tick that meant "we did not look".

How to use this list

  • Do not treat it as your whole manual audit. It is the floor: the criteria you definitely cannot delegate to a scanner. Several of the 23 a tool does check still need a person to confirm the result.
  • Start with the four you will fail. In most audits the first failures are 2.4.7 focus visible, 2.4.3 focus order, 3.3.1 error identification and 1.4.11 non-text contrast. They cost little to fix and they are invisible to every scanner.
  • Run the cross-page ones once per release, not per page. Consistent navigation, consistent identification and consistent help only mean anything across a set.
  • Write down what you did not check. That single habit is the difference between an audit and a reassurance. Our website QA checklist marks which checks a tool can decide and which need a person, and European Accessibility Act testing covers what the law now expects of that record.
  • Work through all 55, not only the 32. Our WCAG 2.2 quick check is a free, in-browser checklist covering every Level A and AA criterion, grouped by principle, with this page's 23-of-55 split marked against each one as you go.

Method

  • The tool side. axe-core 4.10.3, loaded in a real browser, queried with axe.getRules(["wcag2a", "wcag2aa", "wcag21a", "wcag21aa", "wcag22aa"]). Each rule carries tags of the form wcag111, which map to success criterion numbers. 69 rules, carrying tags for 24 distinct criteria, one of which (2.1.3 Keyboard, No Exception) is Level AAA and so falls outside this comparison, leaving 23 at A and AA.
  • The standard side. The success criteria listed in the W3C WCAG 2.2 Recommendation, read from the Recommendation itself: 31 at Level A and 24 at Level AA, 55 in total. Level AAA is out of scope, as it is for almost every legal requirement.
  • The difference is the 32 above. "No rule tagged to it" is the precise claim. A rule may happen to catch something related to a criterion it is not tagged with, and axe's incomplete results flag some of this territory for review rather than ignoring it.
  • Other engines differ at the edges and none of them close this gap, because the gap is not a tooling shortfall. It is what the criteria ask.

Sources, opened and checked 21 September 2026: W3C, Web Content Accessibility Guidelines (WCAG) 2.2; W3C WAI, Selecting Web Accessibility Evaluation Tools; Deque, axe-core 4.10.3; ShipperAG, what automated accessibility testing catches, and what it misses (20 September 2026).

// FAQ

Manual accessibility checks, answered

How much of WCAG can automated testing check?

Less than half of the success criteria, and not the hard half. axe-core 4.10.3, the engine behind most accessibility tooling, ships 69 rules in its WCAG A and AA sets, and those rules carry tags for 23 of the 55 success criteria at Level A and AA in WCAG 2.2. The other 32 have no rule tagged to them at all, because they ask whether something makes sense rather than whether it exists.

Does a clean axe-core run mean a page is accessible?

No. It means no rule that axe can decide failed on the page as it loaded. In our own scan of 410 high-traffic home pages, 81% also returned results axe flagged for human review rather than deciding. A clean automated result is one sentence, not a verdict.

Which WCAG criteria need a person?

The 32 listed on this page, plus a judgement call on several of the 23 a tool does check. The pattern is consistent: a tool can tell you an element exists, has a name, or meets a contrast ratio. It cannot tell you the name is accurate, the order makes sense, the error message helps, or the focus ring is visible against what is behind it.

Can AI do the manual accessibility checks instead?

It can do some of the running and none of the deciding. An agent can tab through a page, trigger every error state and record what happened, which is the tedious part. Whether the resulting experience works for a person using a screen reader is still a judgement, and for a conformance claim it needs someone accountable to make it.

How was this list produced?

By reading axe-core's own rule metadata through axe.getRules() for the wcag2a, wcag2aa, wcag21a, wcag21aa and wcag22aa tag sets, and diffing the WCAG tags those rules carry against the success criteria listed in the W3C's WCAG 2.2 Recommendation. Both halves are reproducible in a few minutes, and the method is written out below.

Can we automate accessibility testing?

Partly, and it is worth knowing exactly how far. A scanner settles the criteria that can be judged from the code: contrast ratios, missing alternative text, a control with no accessible name, a missing language attribute. In axe-core 4.10.3 that is 23 of the 55 WCAG 2.2 success criteria at Level A and AA. The other 32 listed above cannot be automated at all in the sense of reaching a verdict, because they ask whether something makes sense rather than whether it exists. What can be automated is the running: tabbing the page, triggering every error state, comparing across pages. The deciding stays with a person.

How do you do accessibility testing manually?

Work through the 32 criteria above in four passes, because they group naturally. Tab through the whole page with a keyboard and watch the focus indicator: that covers focus order, keyboard traps, focus visibility and focus obscured. Trigger every error state you can reach: that covers error identification, error suggestion and error prevention. Compare the same elements across several pages: that covers consistent navigation, identification and help. Then read the page with a screen reader and judge what is left. Write down what you did not check, which is the part most audits skip.

// free 45-day trial · waitlist open

The judgement is the work. The rest is ours

ShipperAG runs the mechanical half — tabbing the page, triggering the error states, comparing across pages — and reports four states rather than two, so the judgement calls reach you with the evidence attached. Private pilot, for tech and product teams.

We keep what you enter here and the page you joined from, only to contact you about the ShipperAG pilot. No spam.