Short answer. Verify AI-generated code the way you would verify any code: check what it does against the written requirements, not against how the diff reads. Use a different model, person or session to write and run the tests than the one that wrote the code, so its blind spots are not left to check themselves. Keep a record of what was checked and what passed before the release ships. Hand-written, AI-assisted or both get the same checks. ShipperAG runs this for tech and product teams shipping their own product, in a private pilot.
What the research says about AI code quality
AI-generated code usually looks right and runs, so its failures hide where a quick look will not find them: business logic, edge cases, security and accessibility. The published evidence points the same way.
- It compiles, but is it safe? Veracode, an application security vendor, has tested more than 150 models. Its spring 2026 update found syntax correctness above 95%, but only about 55% of generation tasks produced secure code when no security guidance was given.
- Almost right is common. In Stack Overflow's 2025 Developer Survey, the biggest frustration with AI tools, cited by 66% of developers, was "AI solutions that are almost right, but not quite".
- Speed can cost stability. Google's 2025 DORA report found that AI adoption still has a negative relationship with software delivery stability, and that without strong automated testing and fast feedback loops, more change leads to instability.
None of this makes AI-written code bad. It means the volume of change goes up while the signs that something is wrong get quieter. For a software team shipping faster with AI coding tools, that risk lands in production, in front of your users.
Common failure modes in AI-written code
When testing AI-generated code, plan around these failure modes and the checks that catch them.
| Failure mode | What it looks like | Check that catches it |
|---|---|---|
| Plausible but wrong business logic | Discounts stack when they should not, VAT or rounding is off, a rule from the requirements is missing | Test cases written from the requirements and business rules |
| Missing edge cases | Empty basket, expired session, a slow or failed API call, long names, other time zones and languages | Boundary and negative tests, exploratory testing |
| Broken access control | One customer can see another's order by changing an ID; admin routes are reachable | Tests with several roles and accounts, including direct URL and API calls |
| Leaked secrets and risky dependencies | API keys in front-end code or commits; packages that do not exist or are not the ones intended | Secret scanning and dependency review |
| Accessibility gaps | Clickable divs, missing labels, placeholder attributes left in, no visible focus | Automated rules engine plus keyboard and screen reader checks |
| Design system drift | Hard-coded colours and spacing, one-off components, missing states | Visual comparison against design tokens and components |
| Regressions | A change to one component breaks checkout or sign-in elsewhere | End-to-end regression checks on key journeys |
The security rows are well documented. Broken access control is first in the OWASP Top 10:2025, and software supply chain failures are third. GitGuardian's 2026 secrets report found that commits assisted by Claude Code leaked secrets at a rate of 3.2%, against 1.5% for all public GitHub commits, while stressing that people still decide what gets committed. A USENIX Security 2025 study of 576,000 generated code samples found that, on average, at least 5.2% of packages suggested by commercial models, and 21.7% by open-source models, did not exist: an opening for package confusion attacks.
Accessibility needs the same attention. In a CHI 2025 study of 16 developers without accessibility training, the main problems in AI-assisted coding were not asking the AI for accessibility, leaving placeholder attributes in place and being unable to verify compliance.
Why reviewing the diff is not enough
Code review shows how the code reads, not how the product behaves in a browser, with real data, across a whole journey. And AI tools give reviewers more code to read than before.
- Confidence rises faster than quality. In a user study published at ACM CCS 2023, participants with an AI assistant wrote significantly less secure code than those without one, and were more likely to believe their code was secure.
- Verification is the new bottleneck. DORA's March 2026 article on balancing AI tensions notes that the time saved in creating code is frequently re-allocated to auditing and verification.
- Diffs hide context. A reviewer sees the changed lines, not the business rule that the change quietly broke, or the screen three steps later that now fails for keyboard users.
Keep code review. Just do not treat an approved pull request as a tested release.
Why the same model should not write the tests
Asking the tool that wrote the code to write its tests as well feels efficient, but it mostly confirms that the code does what the code does.
- Tests inherit the code's assumptions. A 2024 study of LLM-generated test oracles across 24 open-source Java projects found they tend to capture the program's actual behaviour rather than the expected behaviour. If the code is wrong, the test agrees with it.
- Models favour their own work. A 2024 paper on LLM evaluators describes self-preference: models scoring their own outputs higher than others' that human annotators rate as equal, linked to models recognising their own writing.
- Shared blind spots. If the prompt never mentioned an edge case, neither the code nor its tests will cover it.
AI can still help write tests. The fix is independence: derive tests from the requirements rather than the code, keep them separate from the session that wrote the code, and let deterministic checks, not a model's opinion, decide what passes. We compare scripted, AI-assisted and agentic approaches in agentic QA vs test automation.
A step-by-step QA process for AI-generated code
Here is how to test AI-generated code in eight steps, whichever tool wrote it. The same process works for hand-written code.
- Write down intent first. Before anyone prompts a coding tool, record acceptance criteria, business rules, edge cases and non-functional targets, such as WCAG 2.2 AA and performance budgets. This is what every later check is measured against.
- Test against the requirements, not the code. Derive test cases from that intent. Where the requirements are unclear or contradict each other, ask the product owner instead of letting the code decide.
- Keep checks independent. Use a different person, tool or session to write and run checks from the one that wrote the code, and never let the agent that fixes something approve its own fix.
- Run security checks. Test access with several roles and accounts, scan for secrets, confirm every new dependency exists and is the one you meant, and test input handling on forms and APIs.
- Run accessibility checks. Use an automated rules engine on every build, then keyboard and screen reader checks on key journeys. Our guide to AI accessibility testing covers what tools can and cannot judge.
- Check against the design system. Compare colours, type, spacing and component states with the design tokens, not with the last screenshot.
- Run regression checks on whole journeys. A change made in one place can break a journey elsewhere, so re-test sign-up, checkout and account flows, not only the changed screen. AI regression testing covers how to choose that scope when you have no suite to run.
- Collect evidence before release. Record what was checked, what failed, what was fixed and what was not covered, then get sign-off. Our website QA checklist and UAT sign-off template help.
Take it with you. Download the AI-generated code pre-release checklist (.md), or the whole release readiness kit (.zip). Free under CC BY 4.0. Then score the release with the release readiness scorecard.
Want this checked on a real release of your own product? Start with a free 45-day trial, no card.
Work email only. We keep your email, team size, plan choice and the page you joined from, only to contact you about the ShipperAG pilot. No spam.
What this means for software teams
Whether a prototype was vibe-coded in an afternoon or built over months with Copilot, your users care about one thing: whether the release works. Treat QA for AI-generated code as part of how your team ships, with three habits:
- Budget the testing. Put some of the time AI saves into QA, rather than spending all of it on shipping more.
- Keep the context current. The requirements, design system and business rules are what you test against, so update them when scope changes.
- Share evidence, not assurances. A report showing what was verified, what failed and what was not covered is worth more to stakeholders than "we tested it".
How ShipperAG tests AI-generated code
ShipperAG tests the product the same way however the code was written: by hand, with Cursor, Copilot or Claude Code, or a mix. You connect a preview or staging URL and add context: the requirements, design system, business rules and API docs. The coordinator reads the change and picks the specialists it needs, and they run real checks: a real browser, HTTP and API calls, automated accessibility rules and visual comparison against your design tokens. Its security checks target broken access control, the top risk in the OWASP Top 10.
Language models help read context and write checks, but deterministic checks decide pass or fail, and an independent verifier re-runs every finding in a fresh session. Unclear requirements come back as needs input rather than guesses.
ShipperAG is in a private pilot. See how it works and meet the AI QA specialists.
Sources, checked 19 September 2026: Veracode, Spring 2026 GenAI Code Security Update (24 March 2026); Stack Overflow, 2025 Developer Survey: AI; Google Cloud, Announcing the 2025 DORA Report (24 September 2025); DORA, Balancing AI tensions (10 March 2026); OWASP, Top 10:2025; GitGuardian, The State of Secrets Sprawl 2026 (17 March 2026); Spracklen et al., We Have a Package for You! (USENIX Security 2025); CodeA11y: Making AI Coding Assistants Useful for Accessible Web Development (CHI 2025); Perry et al., Do Users Write More Insecure Code with AI Assistants? (ACM CCS 2023); Konstantinou et al., Do LLMs generate test oracles that capture the actual or the expected program behaviour? (2024); LLM Evaluators Recognize and Favor Their Own Generations (2024); Collins, Word of the Year 2025.