What are AI QA agents?
AI QA agents, also called AI testing agents, are pieces of software that use AI models to test a product on their own. Unlike a test script, an agent can read a requirement, decide how to check it, use tools like a browser or an API client, and interpret what it sees.
The idea sounds simple. The hard part is making an agent's results reliable. Our guide to agentic QA testing covers that in depth. This page focuses on how ShipperAG's agents are organised.
Why specialists instead of one general agent?
Good human QA teams are not made of generalists who test everything at once. Someone thinks about accessibility, someone else about security, someone else about whether the numbers add up. Each person brings a different way of looking for failure.
AI testing agents benefit from the same split:
- Focus. A security agent only thinks about access and data exposure, so it does not get distracted by button colours.
- Right tools for the job. An accessibility agent needs a keyboard and an accessibility tree. A data integrity agent needs to compare what was submitted with what was stored.
- Clear accountability. Every finding says which specialist raised it and what evidence backs it.
- Cost control. A small change does not need every specialist. Only the relevant ones run.
The AI QA specialists
The specialists fall into five groups. Each one checks your product against your own context: your requirements, design system and business rules.
Business3 agents
- Business rules
- Checks prices, limits and rules exactly as your requirements say, including edge values like a threshold of exactly $50.00.
- Product acceptance
- Confirms the feature that was asked for is the one that was built, against its acceptance criteria.
- Analytics
- Makes sure tracking events fire, once, with the right names and data.
Experience8 agents
- User journeys
- Walks key flows like sign-up and checkout from start to finish.
- Personas
- Tests whether your real kinds of users, such as a first-time buyer or an admin, can finish their tasks.
- Accessibility
- Keyboard access, screen reader names, contrast and labels, checked against WCAG.
- Visual design
- Compares screens with your design system at every screen size.
- UX acceptance
- Looks for confusing screens, dead ends and unclear messages.
- Exploratory
- Tries unexpected but realistic actions, like double-submitting or going back mid-flow.
- Localization
- Languages, currencies, dates and time zones.
- Mobile
- Phones and tablets, gestures and app lifecycle.
Engineering7 agents
- API contracts
- Checks every endpoint against its documented contract: fields, types and status codes.
- Code changes
- Reads each change and flags the risky parts, so testing effort goes where it matters.
- Unit and component
- Tests the smallest pieces of code on their own.
- Regression
- Finds what used to work and quietly broke.
- Data integrity
- Makes sure what is saved is exactly what should be saved.
- Migrations
- Checks that old data keeps working after an upgrade.
- Concurrency
- Many users at once, retries and race conditions.
Trust3 agents
- Security
- Checks that nobody can see or change data that is not theirs, such as a viewer editing an order.
- AI features
- Checks that your product's own AI features answer safely and accurately.
- Integrations
- Payments, email and outside services working together.
Operations3 agents
- Performance
- Speed and stability under real and heavy load.
- Notifications
- Emails and alerts reach the right person with the right content.
- Observability
- When something fails, your team can see why and recover.
Want AI QA specialists running on your next release? Start with a free 45-day trial, no card.
Work email only. We keep your email, team size, plan choice and the page you joined from, only to contact you about the ShipperAG pilot. No spam.
The coordinator: choosing the right agents
The coordinator reads what changed and what your context says, then decides which specialists this release needs. A pricing change brings in business rules, data integrity and regression. A new sign-up form brings in user journeys, accessibility, security and analytics.
It works within limits you set, such as time budget or which areas are in scope, so autonomy never means unlimited spend or surprise side effects. We explain those guardrails in autonomous QA testing without test scripts.
The independent verifier: no finding without proof
Every AI model can be confidently wrong. So no specialist's finding goes straight to you. The independent verifier re-runs each finding in a fresh session.
- If an issue reproduces, it is reported as issue found, with screenshots and steps.
- If a pass is backed by a check that actually ran and holds on re-run, it is verified.
- If requirements conflict, it becomes needs input, and a person decides.
- Anything that was not checked is listed as not covered.
A model's opinion is never treated as proof. That one rule does more for trust than any number of extra agents.
What you get from the agents
All findings are written to a hash-chained evidence ledger, recording what was checked on which build, what was found and who approved the release. Product teams can share it with stakeholders and use it for release sign-off. Checks are designed to run in your own CI, so your code and secrets stay with you.