What we found
We scanned all 208 apps listed in Lovable's public gallery on 2 October 2026. 206 returned a page we could measure; 2 were near-empty, with fewer than 40 characters of text, and were left out. None were skipped by robots.txt and none failed to load. Every number below describes those 206 apps.
- 70% of apps (145 of 206) failed at least one WCAG 2.0, 2.1 or 2.2 A or AA check that axe-core can decide on its own. For the web's 410 most-visited sites, the figure was 74%.
- Not one app was missing a language attribute, image alternative text or a page title. On the top sites, the first two each affected 15% of pages.
- 59% of apps (121 of 206) failed colour contrast. On the top sites it was 35%. It was the most common failure in both samples, by a much wider margin here.
- Only 14 distinct rules failed anywhere in the sample, against 35 on the top sites. The median failing app broke 1 rule across 6 elements; the worst broke 4 rules across 75 elements.
- 78% of apps (160 of 206) also returned at least one result axe-core flagged for a person to review instead of deciding itself.
The headline numbers are close: 70% against 74%. The shape underneath them is not. AI-built apps have almost none of the structural failures that still run through the web's biggest sites, and a great deal more of one visual one.
Side by side with the web's top sites
Same scanner, same rules, same browser, same window size. The top-sites column is our September scan of 410 high-traffic home pages.
| Measure | Top 410 sites | 206 AI-built apps |
|---|---|---|
| Failed at least one machine-decidable rule | 74% | 70% |
| Returned at least one check it could not decide | 81% | 78% |
color-contrast | 35% | 59% |
link-name | 24% | 9% |
target-size | 18% | 7% |
image-alt | 15% | 0% |
html-has-lang | 15% | 0% |
button-name | 8% | 10% |
| Distinct rules that failed anywhere | 35 | 14 |
| Median failing page | 2 rules, 3 elements | 1 rule, 6 elements |
Two very different populations: the largest, most heavily resourced sites on the web, and small, recently built apps. That makes the contrast result more striking, not less. A simpler page has fewer chances to fail anything, and these apps fail almost nothing else.
What a machine decided
These are the rules that failed most often across the 206 apps.
| Rule | What it means | Pages | Share |
|---|---|---|---|
color-contrast | Text and its background are too close in colour to read comfortably | 121 | 59% |
button-name | A button has no accessible name | 20 | 10% |
link-name | A link has no text a screen reader can announce | 19 | 9% |
target-size | A control is smaller than the minimum touch target size | 14 | 7% |
scrollable-region-focusable | A scrollable area cannot be reached or scrolled with a keyboard | 7 | 3% |
link-in-text-block | A link is distinguished from surrounding text by colour alone | 5 | 2% |
aria-input-field-name | An ARIA input field has no accessible name | 4 | 2% |
label | A form field has no label | 3 | 1% |
select-name | A select menu has no accessible name | 2 | 1% |
nested-interactive | A control contains another control, which confuses assistive technology | 2 | 1% |
Colour contrast is not one failure among many here; it is most of the result. Of the 201 rule failures we counted, 26 were rated critical and 175 serious by axe-core, and colour contrast accounted for 1248 of the failing elements on its own.
The pattern is consistent with a design default rather than many separate mistakes. One rule failing across several elements is what you would expect from one low-contrast colour, such as a light grey for body or secondary text, used throughout a generated design. That is an inference from the shape of the data; the scan measured the failures, not their cause.
What a machine could not decide
78% of apps returned at least one check axe-core would not settle, on a median of 1 per app and as many as 3. The most common was colour contrast again, on 72% of apps: text over a gradient, an image or an overlay is text whose background the tool cannot read.
So the most common failure in AI-built apps is also the check automation is least able to decide. A tool reports the contrast failures it can measure, and leaves the rest without a verdict.
The clearest number in the study is this one. Of the 61 apps that failed no machine-decidable rule at all, 52 (85%) still had checks nobody decided. A clean automated result on an AI-built app is, more often than not, a result with a gap in it.
What this means before a release
If your team ships with an AI app builder or coding assistant, the structural checks are largely taken care of. The risk has moved somewhere else: into design choices the tool generated and nobody reviewed, and into the checks a scanner cannot settle at all.
- Verified. A check ran and passed, and you can say which check and when.
- Issue found. A check ran and failed, reproducibly. Here, mostly contrast.
- Needs input. The tool reached its limit and a person has to judge — often, here, whether text over a background is readable.
- Not covered. Nothing checked it at all.
A report that shows only pass and fail would call 30% of these apps clean. 85% of those still had an open question. That is the case for reporting four states rather than two, and for checking what an AI tool built against what it was supposed to build. Our guide to testing AI-generated code covers the rest of that, and the 32 WCAG checks automation cannot make lists what a person still has to look at.
Method
- Sample. Every app listed on Lovable's public Discover gallery on 2 October 2026, 208 in all, each at its
lovable.appaddress. These are apps their builders chose to publish to a public gallery, so they may be better finished than the average app made with the tool. - Why one builder. We looked for public galleries of live, deployed apps from other AI app builders. Those we checked either listed templates rather than apps people had shipped, or required signing in. A sample that mixed builders would have been better; we did not find an honest way to build one.
- Pages. One page per app, its home page, following redirects. Nothing was submitted and no other page was requested. Pages with fewer than 40 characters of text were treated as not measurable.
- Browser and rules. Identical to the September study: Chromium through Playwright at 1440 by 900, reduced motion, waiting for the document and then 2.5 seconds; axe-core 4.10.3 restricted to
wcag2a,wcag2aa,wcag21a,wcag21aaandwcag22aa. Best-practice rules excluded. - Politeness. robots.txt read first; one request per app, spaced out; no logins, no forms, no repeat visits.
- Privacy. Aggregates only. No individual app is named here or anywhere else.
What this study is not
It is not a verdict on Lovable, or on AI app builders in general. It describes 206 apps from one builder's public gallery, measured once, on one page each, in one browser at one window size. It is not a WCAG audit of anybody; a conformance claim needs a person, assistive technology and more than one page.
It is also not a like-for-like comparison of AI-made and human-made software. The top sites are large, mature and built by many people over years; these apps are small and new. What the comparison does show is the difference in shape: which failures appear, and which ones do not.
Use the data
The rule-level results are published as a CSV under CC BY 4.0: ai-built-apps-accessibility-study-2026.csv. Credit ShipperAG and link to this page. The scanner is published too, so anyone can repeat the scan exactly: accessibility-study-scanner.mjs, under the MIT licence.
To cite it, copy this:
ShipperAG (2026). In 206 AI-built apps, the web's oldest accessibility bugs are gone. Contrast failures nearly doubled. An axe-core scan of 206 AI-built apps, 2 October 2026. https://shipperag.com/ai-built-apps-accessibility-study/
Sources, all opened and checked 2 October 2026: Lovable, Discover gallery; Deque, axe-core; W3C WAI, WCAG 2.2; ShipperAG, What automated accessibility testing catches, and what it misses (20 September 2026).