Guide · For tech and product teams

AI regression testing for product teams

AI regression testing means using AI to work out what a change could have broken, then running real checks against the new build to find out whether it did. The AI decides the scope; deterministic checks decide pass or fail. It earns its place when a team ships often and has no regression suite to run.

Published · ShipperAG team

What AI regression testing is

A regression is behaviour that used to work and does not any more. AI regression testing is the practice of letting a model read the change, the requirements and the product, decide what that change could plausibly have broken, and then drive real checks against the new build to confirm or clear each suspicion.

The division of labour matters more than the label. The model chooses scope and writes the checks; the checks themselves return facts — an HTTP status, a database row, a rule from an accessibility engine, a rendered page compared against design tokens. A model's opinion that a diff "looks fine" is not a result, and no amount of confident prose makes it one.

Two names for the same idea: AI-powered or AI-driven regression testing usually describes AI bolted onto an existing suite (generating tests, healing selectors). A regression testing AI agent describes something that decides what to check on its own. They solve different problems; the second is the one that helps a team with no suite.

Why regressions are getting harder to catch

Because the number of changes per week went up and the capacity to review them did not. DORA's research puts it plainly: "higher AI adoption is associated with an increase in both software delivery throughput and software delivery instability", and the time saved writing code "is often re-allocated to verification overhead" (DORA, Balancing AI tensions, 10 March 2026).

The shape of the defects has changed too. In the Stack Overflow 2025 Developer Survey, the single biggest frustration with AI tools, named by 66% of respondents, is "AI solutions that are almost right, but not quite", and 45.2% say debugging AI-generated code is more time-consuming. Almost-right code is exactly the kind that passes a review and breaks something two screens away.

None of that is an argument against AI-assisted development. It is an argument for checking the whole blast radius of a change rather than the lines that changed. We keep a sourced collection of AI testing statistics if you need the numbers with their samples and dates attached.

AI regression testing vs a recorded regression suite

A recorded suite proves that the things someone thought of still work. AI regression testing tries to answer a different question: what else could this change have touched? Both are legitimate, and the trade-off is maintenance against certainty.

How the two approaches differ, feature by feature
AspectRecorded regression suiteAI regression testing
What decides the scopeA fixed list, written once by a personThe change itself, read against the requirements
Who maintains itYour team, every time the UI or API movesNothing to maintain; the checks are made per run
Catches the unexpectedNo — only what was written downSometimes, because scope is chosen per change
Repeatable, identical runsYes, and that is its real strengthOnly if findings are re-run and recorded
Cost of a UI changeBroken selectors across many testsRe-read on the next run
Where it failsGoes stale, then gets skippedWeak on intent nobody wrote down

If you already own a healthy suite, keep it: deterministic, identical runs are worth a great deal, and nothing described here replaces them. The choice between writing that suite yourself and delegating the checking is the subject of agentic QA vs test automation, and agentic testing with Playwright covers the same ground from the code side.

Want to see what a change to your own product quietly broke, before a user finds it? Start with a free 45-day trial, no card.

Work email only. We keep your email, team size and the page you joined from, only to contact you about the ShipperAG pilot. No spam.

How an AI regression run works, step by step

The run is only as good as the context it is given. These six steps are what separates a useful run from a model guessing at a website.

  1. Point it at the change. A preview or staging URL for the new build, and what changed: a branch, a ticket, a summary of the work.
  2. Give it the intent. Requirements or a brief, business rules, design tokens, API documentation. Without this, the run can only check that pages load; with it, the run can check that the product does what it is supposed to do.
  3. Let it pick the blast radius. Which journeys, endpoints, data paths and screens the change could reach — including the ones nobody would think to click.
  4. Run real checks. Real browser sessions, HTTP and API calls, an accessibility rules engine, visual comparison against the design system. Not a summary of the code: an execution against the build.
  5. Re-run every finding independently. A failure that only happened once is not a finding. Playwright's own documentation defines a flaky test as one that "failed on the first run, but passed when retried" — a run that cannot tell the two apart wastes your team's afternoon.
  6. Record what was not covered. The silence in a report is the dangerous part. An honest run names what it did not check.

Steps 4 and 5 are where most of the value sits, and they are the ones a model cannot do by talking. Our guide to AI browser testing covers what running in a real browser does and does not prove.

What AI regression testing cannot decide

Four limits, stated plainly, because a regression report that hides them is worse than no report.

  • Intent nobody wrote down. If the requirements are silent or contradictory on what the new behaviour should be, the right answer is to ask, not to guess. A finding of "needs input" is a real outcome.
  • Judgement-based accessibility. A rules engine cannot settle most WCAG success criteria on its own: we went through axe-core's Level A and AA rules and found the criteria automation cannot decide, and what a person has to do for each.
  • Browsers and devices outside the run. ShipperAG's checks run in real Chromium sessions only, including phone and tablet sizes with touch emulation. Firefox and Safari rendering and real physical devices are not covered, and our reports say so rather than implying coverage we do not have.
  • Whether a difference is a regression or the point of the release. A changed price, a moved button and a new empty state all look like regressions to a diffing tool. Only the requirements — or a person — can say which was intended.

How ShipperAG approaches regression

ShipperAG is built around AI QA specialists, picked for each release. On a change, a coordinator reads what moved and picks the specialists it needs within the limits you set — regression, code changes, API contracts, data integrity and migrations are the ones that carry most of the weight here, alongside the journey and visual specialists when the change reaches the interface.

Every finding then goes to an independent verifier, which re-runs it in a fresh session before it can count. Results are sealed in a hash-chained evidence ledger, and come out as a release report a person can approve or share with stakeholders: verified items, issues with screenshots and steps, open questions, and a plain "not covered" section. You can see how the workflow fits together on the homepage, or read an example release report with made-up data.

If your team is testing a lot of AI-written code, the practical checks are collected in how to test AI-generated code.

Sources, all opened and checked 22 September 2026: DORA, Balancing AI tensions (10 March 2026); Stack Overflow, 2025 Developer Survey: AI; Playwright, Retries. ShipperAG is in a private pilot; no availability, pricing or performance claims are made on this page.

// FAQ

AI regression testing, answered

Can regression testing be automated?

Most of it can, and the parts that cannot are worth knowing. Anything with a deterministic answer automates well: HTTP responses, database state after a migration, accessibility rules an engine can decide, visual differences against design tokens, a checkout that either completes or does not. What does not automate is the judgement about whether the new behaviour is correct. A test suite only knows what someone told it to expect, so a change that breaks a rule nobody wrote down passes. That is the gap AI regression testing tries to close: it reads the requirements and the change, then decides what to check, rather than replaying a fixed list.

Can regression testing be done manually?

Yes, and most teams without a suite are doing it now, usually as a click-through of the main journeys before a release. It works until the product gets big enough that nobody can hold the blast radius of a change in their head. The honest problem with manual regression testing is not accuracy but coverage and repeatability: it shrinks under deadline pressure, it is rarely recorded, and two people testing the same release check different things.

What are AI regression testing tools?

They fall into three groups. Test generators write or record test code for you, and you still own and maintain what they produce. Self-healing runners keep existing tests alive when selectors change. Agentic services take the change and the requirements and run the checks themselves, reporting findings rather than a suite. The three fail differently, so the useful question is what you are left maintaining afterwards. Our comparison of AI testing tools sets out where each group fits.

What is AI visual regression testing?

It compares how a page renders before and after a change and flags the differences that matter. The hard part is not taking screenshots but ignoring the noise: animation, fonts loading, dates, and content that legitimately changed. Comparing against design tokens and brand rules, rather than against yesterday's screenshot alone, is what turns a pile of pixel differences into a finding a person can act on.

Is regression testing done in production?

Some of it is, carefully: synthetic monitoring and health checks watch production for behaviour that broke after a deploy. Most regression testing belongs earlier, against a preview or staging build, because that is where a finding can still stop a release. ShipperAG is designed to run against preview and staging builds, and it only tests targets you own or have written permission to test.

// free 45-day trial · waitlist open

Catch what a change quietly broke

ShipperAG checks every release against what it is supposed to do, re-checks what it finds, and hands your team evidence it can sign. It is in a private pilot; joining the waitlist is how a team becomes a product partner.

  • Free 45-day trial, no card
  • Up to 20 release checks on your product
  • Direct line to the founders

Work email only. We keep your email, team size and the page you joined from, only to contact you about the ShipperAG pilot. No spam.