Reference · Last checked 20 September 2026

AI testing statistics with a link to every source

Forty statistics on AI in software development, testing practice, the cost of defects, accessibility and security. Every figure names what it measures, the sample it came from and the year, and links to the report, survey or study that published it. Nothing here is taken from an aggregator, and no ShipperAG data is included.

Published · ShipperAG team · Figures checked 20 September 2026

AI use in software development

AI tools are now standard in development, but the headline numbers differ by survey, because each one asks a different population a different question. Read the sample before comparing them.

  • 84% of respondents are using or planning to use AI tools in their development process, up from 76% in 2024 (Stack Overflow Developer Survey 2025, more than 49,000 respondents in 177 countries).
  • 51% of professional developers use AI tools daily; across all respondents, 47.1% do (Stack Overflow Developer Survey 2025).
  • 90% of survey respondents report using AI at work, a 14.1% increase on the same metric a year earlier (DORA, State of AI-assisted Software Development 2025, 4,867 technology professionals, surveyed 13 June to 21 July 2025).
  • Two hours is the median time AI users spent interacting with AI on their most recent workday, about a quarter of an eight-hour day (DORA, State of AI-assisted Software Development 2025, AI adoption and use chapter).
  • 85% of developers regularly use AI tools for coding and development, and 62% rely on at least one AI coding assistant, agent or code editor (JetBrains State of Developer Ecosystem 2025, 24,534 developers in 194 countries, surveyed April to June 2025).
  • Nearly 80% of new developers on GitHub use GitHub Copilot within their first week, out of more than 36 million who joined in the year to 2025 (GitHub Octoverse 2025).
  • 50% of code characters at Google are completed by AI-based suggestions, at a 37% acceptance rate, measured in Google's internal tooling (Google Research, June 2024).
  • 71% of respondents who write code use AI to assist them in doing so, the single most common use of AI; among respondents whose work involves each task, 66% use AI to modify existing code and 49% to analyse requirements (DORA 2025, AI adoption and use chapter).

Trust and quality of AI-written code

Adoption has run ahead of trust. In both of the largest 2025 developer surveys, most respondents use AI daily and a minority say they trust what it produces.

  • 46% of developers distrust the accuracy of AI tools and 33% trust it; only 3% highly trust the output (Stack Overflow Developer Survey 2025).
  • 66% name "AI solutions that are almost right, but not quite" as their biggest frustration, the single largest answer (Stack Overflow Developer Survey 2025).
  • 45% say debugging AI-generated code is more time-consuming, the second-biggest frustration (Stack Overflow Developer Survey 2025, 45.2%).
  • Favourable sentiment towards AI tools fell to 60% in 2025, from more than 70% in both 2023 and 2024 (Stack Overflow Developer Survey 2025).
  • 30% report little or no trust in the quality of AI-generated output ("a little" 23%, "not at all" 7%), while 24% report "a great deal" or "a lot" (DORA 2025, AI adoption and use chapter).
  • More than 80% of respondents perceive that AI increased their productivity, and fewer than 10% perceive any decrease (DORA 2025, AI adoption and use chapter).
  • Experienced open-source developers took 19% longer to finish issues when allowed to use early-2025 AI tools, having forecast a 24% speed-up, and still believed afterwards that they had been sped up by 20% (METR randomised controlled trial, July 2025, 16 developers working on their own repositories).
  • Only 55% of AI code generation tasks produced secure code, so 45% introduced a known flaw, while syntax correctness exceeded 95% (Veracode, Spring 2026 GenAI Code Security Update, 80 coding tasks across four languages, more than 150 models tested to date, no security guidance given in the prompt).

Those last two points are the reason this page exists: speed and confidence rise faster than evidence does. Our guide to testing AI-generated code sets out the checks that catch what a diff review misses.

Testing practice and automation

Quality engineering is adopting AI much more slowly than development is. Most organisations are experimenting; few have scaled it.

  • 43% of organisations are experimenting with generative AI in QA, but only 15% have scaled it enterprise-wide (World Quality Report 2025-26, Capgemini and Sogeti; the public summary does not state the sample size).
  • 60% of organisations struggle with secure, scalable test data and 58% cite challenges in adopting AI-powered tools (World Quality Report 2025-26).
  • Synthetic data in testing rose from 14% in 2024 to 25% in 2025, and is the top-ranked generative AI use case in QA (World Quality Report 2025-26).
  • Generative AI is the top-ranked skill for quality engineers (63%), ahead of core quality engineering skills (60%) (World Quality Report 2025-26).
  • 72.6% of developers who use Copilot code review said it improved their effectiveness, from in-depth interviews about the code review process (GitHub Octoverse 2025).
  • More than 1,000,000 ISTQB Certified Tester certificates have been awarded worldwide since the scheme began in 1998, a milestone announced on 6 May 2025 (ISTQB).

Two DORA findings belong beside these. Higher AI adoption is associated with an increase in both software delivery throughput and software delivery instability, and a March 2026 thematic analysis of 1,110 open-ended responses from Google software engineers found verification overhead and hallucinations reported across all ten AI use cases studied, testing included (DORA, Balancing AI tensions, 10 March 2026).

Cost and impact of defects

The public cost estimates are large, old and rarely updated: the most recent broad figure for the United States is from 2022, and the most quoted testing-specific figure is from 2002. Use them with their dates attached.

  • The cost of poor software quality in the US in 2022 was at least $2.41 trillion, with accumulated software technical debt of about $1.52 trillion (CISQ, Cost of Poor Software Quality in the US: A 2022 Report).
  • An inadequate infrastructure for software testing cost the US economy $59.5 billion a year, of which $22.2 billion could be avoided by feasible improvements to testing infrastructure (NIST Planning Report 02-3, 2002, based on developer and user surveys).
  • 82% of organisations carry security debt and 60% carry critical security debt, up from 74% and 50% a year earlier; security debt means flaws left unresolved for more than a year (Veracode, 2026 State of Software Security, published 24 February 2026).
  • Average fix time for critical-severity vulnerabilities fell from 37 to 26 days on GitHub in 2025, a 30% improvement, while 26% fewer repositories received critical alerts (GitHub Octoverse 2025).

Want evidence like this for your own release, not just the industry's? Start with a free 45-day trial, no card.

Work email only. We keep your email, team size, plan choice and the page you joined from, only to contact you about the ShipperAG pilot. No spam.

Accessibility and compliance

Automated checks alone fail on almost every popular home page, and the same six problems account for nearly all the errors that tools can detect. Automation finds a majority of issues by volume, not by success criterion.

  • 95.9% of the top 1,000,000 home pages had detected WCAG 2 failures in February 2026, up from 94.8% in 2025 and reversing six years of small improvements (WebAIM Million 2026).
  • Home pages averaged 56.1 detected accessibility errors, 10.1% more than the 51 errors per page found in 2025 (WebAIM Million 2026).
  • Six failure types account for 96% of all detected errors (WebAIM Million 2026), and they have been the same six for seven years:
Percentage of the top 1,000,000 home pages with each failure, February 2026 (WebAIM Million)
Failure type20262025
Low contrast text83.9%79.1%
Missing alternative text for images53.1%55.5%
Missing form input labels51%48.2%
Empty links46.3%45.4%
Empty buttons30.6%29.6%
Missing document language13.5%15.8%
  • Home pages that used ARIA averaged 59.1 errors, against 42 on pages without it; 82.7% of home pages used ARIA, up from 79.4% in 2025 (WebAIM Million 2026).
  • 57.38% of accessibility issues were found by automated tests in a sample of more than 13,000 pages and page states and nearly 300,000 issues from first-time audits; the same analysis found automatically detectable issues for 16 of the 50 WCAG 2.1 AA success criteria (Deque, The Automated Accessibility Coverage Report, page published November 2025).
  • 71.6% of screen reader users navigate through the headings on a page first when looking for information on a long page (WebAIM Screen Reader User Survey #10, 1,539 valid responses, December 2023 to January 2024).
  • An estimated 1.3 billion people, or 16% of the world's population, experience significant disability (World Health Organization, fact sheet dated 7 March 2023).

For context on what those failures are measured against: WCAG 2.2 became a W3C Recommendation on 5 October 2023 and adds nine success criteria to WCAG 2.1 (W3C Web Accessibility Initiative). Our guide to pre-launch website QA covers where these checks belong in a release.

Security

Access control is the top application security risk, and AI-assisted development is adding two new problems at scale: leaked secrets and package names that do not exist.

  • Broken access control is the number one risk in the OWASP Top 10:2025, and on average 3.73% of applications tested had one or more of its 40 CWEs, from data donated on more than 2.8 million applications; software supply chain failures entered the list at number three (OWASP Top 10:2025).
  • AI-generated code passed security checks 15% of the time for cross-site scripting and 13% for log injection, against 82% for SQL injection and 86% for insecure cryptographic algorithms; by language, Java scored 29% and Python 62% (Veracode, Spring 2026 GenAI Code Security Update).
  • 28.65 million new hardcoded secrets were added to public GitHub commits in 2025, a 34% rise year on year and the largest single-year jump GitGuardian has recorded (GitGuardian, The State of Secrets Sprawl 2026).
  • Commits assisted by Claude Code showed a 3.2% secret-leak rate, against a 1.5% baseline across all public GitHub commits; GitGuardian stresses that developers still decide what is committed (GitGuardian, 2026).
  • 64% of credentials confirmed valid in 2022 were still valid when retested in January 2026, so leaked secrets are rarely rotated (GitGuardian, 2026).
  • At least 5.2% of packages recommended by commercial models, and 21.7% by open-source models, did not exist, including 205,474 unique hallucinated package names, an opening for package confusion attacks (Spracklen et al., USENIX Security 2025, 576,000 generated code samples from 16 LLMs).
  • Broken access control overtook injection as the most common CodeQL alert, flagged in more than 151,000 repositories, up 172% year on year, which GitHub attributes partly to misconfigured CI/CD permissions and AI-generated scaffolds that skip authorisation checks (GitHub Octoverse 2025).

How to cite this page and how we maintain it

Cite it as: ShipperAG, "AI testing statistics", shipperag.com/ai-testing-statistics/, last checked 20 September 2026. Where you can, cite the original report as well; every figure above links to it.

Last checked: 20 September 2026. We re-check every figure each quarter, replace superseded editions with the newest one and remove anything we cannot still see at its source. Rules we follow on this page:

  • Every number is read on the page that published it: the report, survey, standards body, study or the organisation's own write-up. No aggregators, no second-hand citations.
  • Every number carries what it measures, the population or sample and the year, because a percentage without a denominator is not evidence.
  • Figures we cannot verify at a primary source are left out. The often-quoted claim that a defect costs 100 times more to fix in production than in design is one of them: it is usually attributed to an "IBM Systems Sciences Institute" study that no one has been able to produce, as The Register reported in 2021. It is not on this page.

If you spot a figure that has moved, changed or disappeared at its source, tell us and we will correct it. For the tooling side of the same question, see AI testing tools compared.

Sources, all opened and checked 20 September 2026: Stack Overflow, 2025 Developer Survey: AI and 2025 Developer Survey; DORA, State of AI-assisted Software Development 2025 (report page) and Balancing AI tensions (10 March 2026); Google Cloud, Announcing the 2025 DORA report; JetBrains, The State of Developer Ecosystem 2025; GitHub, Octoverse 2025; Google Research, AI in software engineering at Google (6 June 2024); METR, Measuring the impact of early-2025 AI on experienced open-source developer productivity; Veracode, Spring 2026 GenAI Code Security Update and 2026 State of Software Security; Capgemini and Sogeti, World Quality Report 2025-26; ISTQB, 1,000,000+ certificates; CISQ, Cost of Poor Software Quality in the US: A 2022 Report; NIST, Planning Report 02-3 (2002); WebAIM, The WebAIM Million 2026 and Screen Reader User Survey #10; Deque, The Automated Accessibility Coverage Report; World Health Organization, Disability fact sheet (7 March 2023); W3C WAI, What's new in WCAG 2.2; OWASP, Top 10:2025 introduction; GitGuardian, The State of Secrets Sprawl 2026; Spracklen et al., We Have a Package for You! (USENIX Security 2025); The Register, on the missing 100x study (22 July 2021).

FAQ

AI testing statistics, answered

How much of code is now written by AI?

There is no reliable industry-wide figure, only measurements from single organisations and surveys of developers. Google Research reported in June 2024 that AI-based suggestions completed 50% of code characters in its internal tools, at an acceptance rate of 37%. In DORA's 2025 report, 71% of respondents who write code use AI to help them do it. Stack Overflow's 2025 survey found 84% of respondents are using or planning to use AI tools. Treat any single percentage of "AI-written code" as a claim about one company's tooling, not an industry average.

Do developers trust AI-generated code?

Most do not trust it fully. In the Stack Overflow Developer Survey 2025, 46% of developers said they distrust the accuracy of AI tools against 33% who trust it, and only 3% said they highly trust the output. In DORA's 2025 report, 30% reported little or no trust in AI-generated code, while 24% reported a great deal or a lot of trust.

Is AI-generated code less secure?

Veracode's Spring 2026 update found that across all models and tasks only 55% of generation tasks produced secure code, so 45% introduced a known flaw, while syntax correctness exceeded 95%. The weakest areas were cross-site scripting, with a 15% security pass rate, and log injection at 13%. A USENIX Security 2025 study of 576,000 generated samples found that on average at least 5.2% of packages recommended by commercial models did not exist, rising to 21.7% for open-source models.

Does AI make developers faster?

The evidence is mixed. In DORA's 2025 survey, more than 80% of respondents perceived that AI had increased their productivity and fewer than 10% perceived a decrease. In a 2025 randomised controlled trial by METR, 16 experienced open-source developers took 19% longer to complete issues when allowed to use early-2025 AI tools, although they had expected a 24% speed-up and still believed afterwards that they had been 20% faster.

What percentage of websites fail accessibility testing?

In the WebAIM Million analysis of the top 1,000,000 home pages in February 2026, 95.9% had detected WCAG 2 failures, up from 94.8% in 2025, with an average of 56.1 errors per page. Because only automatically detectable failures were counted, WebAIM notes the true conformance rate is certainly lower than 4.1%.

How much does poor software quality cost?

The most recent public estimate is CISQ's 2022 report, which put the cost of poor software quality in the US at at least $2.41 trillion, with accumulated software technical debt of about $1.52 trillion. The older NIST planning report of 2002 estimated the annual cost of an inadequate software testing infrastructure to the US economy at $59.5 billion, of which $22.2 billion could be avoided through better testing infrastructure.

Free 45-day trial · waitlist open

Numbers are easy. Evidence is the work

ShipperAG checks every release against your requirements, design system and business rules, re-checks every finding in a fresh session, and hands your team a report it can sign or share. It is in a private pilot: joining the waitlist is how a team becomes a product partner.

  • Free 45-day trial, no card
  • Up to 20 release checks on your product
  • Direct line to the founders

Work email only. We keep your email, team size, plan choice and the page you joined from, only to contact you about the ShipperAG pilot. No spam.