Guide · AI testing agents

AI QA agents: specialists picked for each release

ShipperAG does not use one AI agent that tries to test everything. It draws on a roster of AI QA agents, each an expert in one way software goes wrong, plus a coordinator that brings in the ones a release needs and a verifier that checks their work.

Published · ShipperAG team

What are AI QA agents?

AI QA agents, also called AI testing agents, are pieces of software that use AI models to test a product on their own. Unlike a test script, an agent can read a requirement, decide how to check it, use tools like a browser or an API client, and interpret what it sees.

The idea sounds simple. The hard part is making an agent's results reliable. Our guide to agentic QA testing covers that in depth. This page focuses on how ShipperAG's agents are organised.

Why specialists instead of one general agent?

Good human QA teams are not made of generalists who test everything at once. Someone thinks about accessibility, someone else about security, someone else about whether the numbers add up. Each person brings a different way of looking for failure.

AI testing agents benefit from the same split:

  • Focus. A security agent only thinks about access and data exposure, so it does not get distracted by button colours.
  • Right tools for the job. An accessibility agent needs a keyboard and an accessibility tree. A data integrity agent needs to compare what was submitted with what was stored.
  • Clear accountability. Every finding says which specialist raised it and what evidence backs it.
  • Cost control. A small change does not need every specialist. Only the relevant ones run.

The AI QA specialists

The specialists fall into five groups. Each one checks your product against your own context: your requirements, design system and business rules.

Business3 agents

Business rules
Checks prices, limits and rules exactly as your requirements say, including edge values like a threshold of exactly $50.00.
Product acceptance
Confirms the feature that was asked for is the one that was built, against its acceptance criteria.
Analytics
Makes sure tracking events fire, once, with the right names and data.

Experience8 agents

User journeys
Walks key flows like sign-up and checkout from start to finish.
Personas
Tests whether your real kinds of users, such as a first-time buyer or an admin, can finish their tasks.
Accessibility
Keyboard access, screen reader names, contrast and labels, checked against WCAG.
Visual design
Compares screens with your design system at every screen size.
UX acceptance
Looks for confusing screens, dead ends and unclear messages.
Exploratory
Tries unexpected but realistic actions, like double-submitting or going back mid-flow.
Localization
Languages, currencies, dates and time zones.
Mobile
Phones and tablets, gestures and app lifecycle.

Engineering7 agents

API contracts
Checks every endpoint against its documented contract: fields, types and status codes.
Code changes
Reads each change and flags the risky parts, so testing effort goes where it matters.
Unit and component
Tests the smallest pieces of code on their own.
Regression
Finds what used to work and quietly broke.
Data integrity
Makes sure what is saved is exactly what should be saved.
Migrations
Checks that old data keeps working after an upgrade.
Concurrency
Many users at once, retries and race conditions.

Trust3 agents

Security
Checks that nobody can see or change data that is not theirs, such as a viewer editing an order.
AI features
Checks that your product's own AI features answer safely and accurately.
Integrations
Payments, email and outside services working together.

Operations3 agents

Performance
Speed and stability under real and heavy load.
Notifications
Emails and alerts reach the right person with the right content.
Observability
When something fails, your team can see why and recover.

Want AI QA specialists running on your next release? Start with a free 45-day trial, no card.

Work email only. We keep your email, team size, plan choice and the page you joined from, only to contact you about the ShipperAG pilot. No spam.

The coordinator: choosing the right agents

The coordinator reads what changed and what your context says, then decides which specialists this release needs. A pricing change brings in business rules, data integrity and regression. A new sign-up form brings in user journeys, accessibility, security and analytics.

It works within limits you set, such as time budget or which areas are in scope, so autonomy never means unlimited spend or surprise side effects. We explain those guardrails in autonomous QA testing without test scripts.

The independent verifier: no finding without proof

Every AI model can be confidently wrong. So no specialist's finding goes straight to you. The independent verifier re-runs each finding in a fresh session.

  • If an issue reproduces, it is reported as issue found, with screenshots and steps.
  • If a pass is backed by a check that actually ran and holds on re-run, it is verified.
  • If requirements conflict, it becomes needs input, and a person decides.
  • Anything that was not checked is listed as not covered.

A model's opinion is never treated as proof. That one rule does more for trust than any number of extra agents.

What you get from the agents

All findings are written to a hash-chained evidence ledger, recording what was checked on which build, what was found and who approved the release. Product teams can share it with stakeholders and use it for release sign-off. Checks are designed to run in your own CI, so your code and secrets stay with you.

FAQ

AI QA agents, answered

What is an AI QA agent?

An AI QA agent is software that uses AI models to test a product on its own. It can read requirements, decide how to check them, use tools like a browser or API client, and interpret the results.

How many AI testing agents does ShipperAG use?

ShipperAG doesn't run a fixed number of specialists. It draws from specialists across five groups (business, experience, engineering, trust and operations), plus a coordinator and an independent verifier, and only the specialists a change needs are run.

How do you stop AI QA agents from making things up?

An independent verifier re-runs every finding in a fresh session before it is reported. Passes must come from checks that actually ran, and unclear or unchecked areas are shown as needs input or not covered.

Do the AI QA agents need access to my source code?

Most checks run against a preview or staging build, using your requirements, design system and business rules as context. Checks are designed to run in your own CI, next to your code.

Can AI replace QA engineers?

No, and the question hides the useful one. AI takes over the repetitive part: running the same checks on every release, re-running them after a fix, and writing down what was covered. It cannot decide what the product should do, judge whether a screen makes sense to a first-time user, weigh a risk against a deadline, or sign off a release. Those are the parts QA engineers are paid for, and the automatic part has always been the part they most want to hand over.

Free 45-day trial · waitlist open

Put AI QA agents on your next release

ShipperAG is in a private pilot. Join the waitlist to become a product partner and try it on a real product.

  • Free 45-day trial, no card
  • Up to 20 release checks on your product
  • Direct line to the founders

Work email only. We keep your email, team size, plan choice and the page you joined from, only to contact you about the ShipperAG pilot. No spam.