Guide · AI-assisted development for software teams

How to test AI-generated code before it reaches your users

Test AI-generated code the way you would test any code you did not write: against the requirements, not against the code. Write down what it should do first, run independent checks that the coding tool did not write, add security, accessibility and regression checks, and keep the evidence before release.

Published · ShipperAG team

Short answer. Verify AI-generated code the way you would verify any code: check what it does against the written requirements, not against how the diff reads. Use a different model, person or session to write and run the tests than the one that wrote the code, so its blind spots are not left to check themselves. Keep a record of what was checked and what passed before the release ships. Hand-written, AI-assisted or both get the same checks. ShipperAG runs this for tech and product teams shipping their own product, in a private pilot.

What the research says about AI code quality

AI-generated code usually looks right and runs, so its failures hide where a quick look will not find them: business logic, edge cases, security and accessibility. The published evidence points the same way.

  • It compiles, but is it safe? Veracode, an application security vendor, has tested more than 150 models. Its spring 2026 update found syntax correctness above 95%, but only about 55% of generation tasks produced secure code when no security guidance was given.
  • Almost right is common. In Stack Overflow's 2025 Developer Survey, the biggest frustration with AI tools, cited by 66% of developers, was "AI solutions that are almost right, but not quite".
  • Speed can cost stability. Google's 2025 DORA report found that AI adoption still has a negative relationship with software delivery stability, and that without strong automated testing and fast feedback loops, more change leads to instability.

None of this makes AI-written code bad. It means the volume of change goes up while the signs that something is wrong get quieter. For a software team shipping faster with AI coding tools, that risk lands in production, in front of your users.

Common failure modes in AI-written code

When testing AI-generated code, plan around these failure modes and the checks that catch them.

What goes wrong, and what catches it
Failure modeWhat it looks likeCheck that catches it
Plausible but wrong business logicDiscounts stack when they should not, VAT or rounding is off, a rule from the requirements is missingTest cases written from the requirements and business rules
Missing edge casesEmpty basket, expired session, a slow or failed API call, long names, other time zones and languagesBoundary and negative tests, exploratory testing
Broken access controlOne customer can see another's order by changing an ID; admin routes are reachableTests with several roles and accounts, including direct URL and API calls
Leaked secrets and risky dependenciesAPI keys in front-end code or commits; packages that do not exist or are not the ones intendedSecret scanning and dependency review
Accessibility gapsClickable divs, missing labels, placeholder attributes left in, no visible focusAutomated rules engine plus keyboard and screen reader checks
Design system driftHard-coded colours and spacing, one-off components, missing statesVisual comparison against design tokens and components
RegressionsA change to one component breaks checkout or sign-in elsewhereEnd-to-end regression checks on key journeys

The security rows are well documented. Broken access control is first in the OWASP Top 10:2025, and software supply chain failures are third. GitGuardian's 2026 secrets report found that commits assisted by Claude Code leaked secrets at a rate of 3.2%, against 1.5% for all public GitHub commits, while stressing that people still decide what gets committed. A USENIX Security 2025 study of 576,000 generated code samples found that, on average, at least 5.2% of packages suggested by commercial models, and 21.7% by open-source models, did not exist: an opening for package confusion attacks.

Accessibility needs the same attention. In a CHI 2025 study of 16 developers without accessibility training, the main problems in AI-assisted coding were not asking the AI for accessibility, leaving placeholder attributes in place and being unable to verify compliance.

Why reviewing the diff is not enough

Code review shows how the code reads, not how the product behaves in a browser, with real data, across a whole journey. And AI tools give reviewers more code to read than before.

  • Confidence rises faster than quality. In a user study published at ACM CCS 2023, participants with an AI assistant wrote significantly less secure code than those without one, and were more likely to believe their code was secure.
  • Verification is the new bottleneck. DORA's March 2026 article on balancing AI tensions notes that the time saved in creating code is frequently re-allocated to auditing and verification.
  • Diffs hide context. A reviewer sees the changed lines, not the business rule that the change quietly broke, or the screen three steps later that now fails for keyboard users.

Keep code review. Just do not treat an approved pull request as a tested release.

Why the same model should not write the tests

Asking the tool that wrote the code to write its tests as well feels efficient, but it mostly confirms that the code does what the code does.

  • Tests inherit the code's assumptions. A 2024 study of LLM-generated test oracles across 24 open-source Java projects found they tend to capture the program's actual behaviour rather than the expected behaviour. If the code is wrong, the test agrees with it.
  • Models favour their own work. A 2024 paper on LLM evaluators describes self-preference: models scoring their own outputs higher than others' that human annotators rate as equal, linked to models recognising their own writing.
  • Shared blind spots. If the prompt never mentioned an edge case, neither the code nor its tests will cover it.

AI can still help write tests. The fix is independence: derive tests from the requirements rather than the code, keep them separate from the session that wrote the code, and let deterministic checks, not a model's opinion, decide what passes. We compare scripted, AI-assisted and agentic approaches in agentic QA vs test automation.

A step-by-step QA process for AI-generated code

Here is how to test AI-generated code in eight steps, whichever tool wrote it. The same process works for hand-written code.

  1. Write down intent first. Before anyone prompts a coding tool, record acceptance criteria, business rules, edge cases and non-functional targets, such as WCAG 2.2 AA and performance budgets. This is what every later check is measured against.
  2. Test against the requirements, not the code. Derive test cases from that intent. Where the requirements are unclear or contradict each other, ask the product owner instead of letting the code decide.
  3. Keep checks independent. Use a different person, tool or session to write and run checks from the one that wrote the code, and never let the agent that fixes something approve its own fix.
  4. Run security checks. Test access with several roles and accounts, scan for secrets, confirm every new dependency exists and is the one you meant, and test input handling on forms and APIs.
  5. Run accessibility checks. Use an automated rules engine on every build, then keyboard and screen reader checks on key journeys. Our guide to AI accessibility testing covers what tools can and cannot judge.
  6. Check against the design system. Compare colours, type, spacing and component states with the design tokens, not with the last screenshot.
  7. Run regression checks on whole journeys. A change made in one place can break a journey elsewhere, so re-test sign-up, checkout and account flows, not only the changed screen. AI regression testing covers how to choose that scope when you have no suite to run.
  8. Collect evidence before release. Record what was checked, what failed, what was fixed and what was not covered, then get sign-off. Our website QA checklist and UAT sign-off template help.

Take it with you. Download the AI-generated code pre-release checklist (.md), or the whole release readiness kit (.zip). Free under CC BY 4.0. Then score the release with the release readiness scorecard.

Want this checked on a real release of your own product? Start with a free 45-day trial, no card.

Work email only. We keep your email, team size, plan choice and the page you joined from, only to contact you about the ShipperAG pilot. No spam.

What this means for software teams

Whether a prototype was vibe-coded in an afternoon or built over months with Copilot, your users care about one thing: whether the release works. Treat QA for AI-generated code as part of how your team ships, with three habits:

  • Budget the testing. Put some of the time AI saves into QA, rather than spending all of it on shipping more.
  • Keep the context current. The requirements, design system and business rules are what you test against, so update them when scope changes.
  • Share evidence, not assurances. A report showing what was verified, what failed and what was not covered is worth more to stakeholders than "we tested it".

How ShipperAG tests AI-generated code

ShipperAG tests the product the same way however the code was written: by hand, with Cursor, Copilot or Claude Code, or a mix. You connect a preview or staging URL and add context: the requirements, design system, business rules and API docs. The coordinator reads the change and picks the specialists it needs, and they run real checks: a real browser, HTTP and API calls, automated accessibility rules and visual comparison against your design tokens. Its security checks target broken access control, the top risk in the OWASP Top 10.

Language models help read context and write checks, but deterministic checks decide pass or fail, and an independent verifier re-runs every finding in a fresh session. Unclear requirements come back as needs input rather than guesses.

ShipperAG is in a private pilot. See how it works and meet the AI QA specialists.

Sources, checked 19 September 2026: Veracode, Spring 2026 GenAI Code Security Update (24 March 2026); Stack Overflow, 2025 Developer Survey: AI; Google Cloud, Announcing the 2025 DORA Report (24 September 2025); DORA, Balancing AI tensions (10 March 2026); OWASP, Top 10:2025; GitGuardian, The State of Secrets Sprawl 2026 (17 March 2026); Spracklen et al., We Have a Package for You! (USENIX Security 2025); CodeA11y: Making AI Coding Assistants Useful for Accessible Web Development (CHI 2025); Perry et al., Do Users Write More Insecure Code with AI Assistants? (ACM CCS 2023); Konstantinou et al., Do LLMs generate test oracles that capture the actual or the expected program behaviour? (2024); LLM Evaluators Recognize and Favor Their Own Generations (2024); Collins, Word of the Year 2025.

FAQ

Testing AI-generated code, answered

Is AI-generated code less secure than human-written code?

It can be. In a user study published at ACM CCS 2023, people using an AI assistant wrote significantly less secure code and were more confident in it. Veracode's 2026 tests found that, without security guidance, about 45% of AI generation tasks introduced a known security flaw. Test all code the same way.

Can I use AI to write tests for AI-generated code?

Yes, with care. Base the tests on the requirements rather than the code, keep them separate from the session that wrote the code, and review them. A 2024 study found that LLM-generated test oracles tend to capture what the code does, not what it should do.

What is vibe coding?

Collins Dictionary, which named it its Word of the Year 2025, defines vibe coding as the use of artificial intelligence prompted by natural language to write computer code. Testing a vibe-coded app is no different from testing any other: check the product against what it should do.

Do I need different tests for code from Cursor, Copilot or Claude Code?

No. The tool does not change what the product must do, so test against the requirements, design system and business rules in every case. With AI-assisted code, give extra attention to access control, secrets and new dependencies.

Who should approve AI-generated code for release?

A person, with evidence. Automated checks and AI agents can gather the evidence, but the tool that wrote or fixed the code should never approve its own work. The release owner, such as an engineering lead or product owner, signs off on a report of what was verified.

How do you verify AI-generated code before using it?

Read it, then prove it. Read the change against what the feature is supposed to do, not against how the code reads, because a model writes plausible code for the wrong requirement as easily as for the right one. Then run checks that can fail: the behaviour the requirement describes, the edge cases the model did not consider, the access rules around the data it touches, and the areas the change could break. A second pass by whoever wrote the code, human or model, is not verification.

What is AI code verification?

AI code verification is checking that code written by an AI tool does what the requirement asked for, rather than checking that it compiles and looks plausible. It has three parts: read the change against the requirement rather than against how the code reads; run checks that can actually fail, covering the behaviour specified, the edge cases the model did not consider and the access rules around the data it touches; and make sure whatever verifies is not the same model that generated. A second pass by the author, human or model, is review, not verification.

Free 45-day trial · waitlist open

Test the product, not the prompt

ShipperAG checks every release against your requirements, design system and business rules, however the code was written, and re-checks every finding before it reaches your report. It is in a private pilot: join the waitlist to become a product partner.

  • Free 45-day trial, no card
  • Up to 20 release checks on your product
  • Direct line to the founders

Work email only. We keep your email, team size, plan choice and the page you joined from, only to contact you about the ShipperAG pilot. No spam.