Reference · QA and AI testing terms

QA and AI testing glossary: 59 terms, defined plainly

This glossary defines the words engineering and product teams use when they talk about testing a release: the classic QA vocabulary, the newer agentic and AI terms, and the standards and tools people name in passing. Each entry answers "what is this" in one sentence, then adds the part that usually matters.

Published · ShipperAG team · Sources checked 20 September 2026

Testing types

These are the kinds of testing a team can run, grouped by what each one is trying to prove. Most products use several at once, at different points in the delivery cycle.

Unit testing

Unit testing checks one small piece of code, such as a function or a component, in isolation from the rest of the system. Dependencies are replaced with stubs, so the test is fast and the failure points at a single place. ISTQB calls this component testing.

Integration testing

Integration testing checks that separate components or systems work correctly together, such as your code and its database, or your service and a payment provider. It finds faults in the joins between parts, like wrong formats or missing fields, which unit tests never see because every part passes alone.

End-to-end testing

End-to-end testing exercises a complete user journey through the real system, from the interface down to the database and back. A test signs in, adds an item, pays and reads the confirmation. It is the slowest and most realistic functional test, and the one whose failures a user would actually notice.

Regression testing

Regression testing re-checks behaviour that already worked, to find defects a change has introduced or uncovered. It answers one question: did this release break something that was fine yesterday? Teams run it on every build, because regressions usually appear in code nobody touched. ISTQB defines it as change-related testing.

Smoke testing

A smoke test is a short set of checks that proves a build is stable enough to be worth testing properly: the app starts, the main pages load, sign-in works. If it fails, deeper testing stops until someone fixes the build, which saves a day of chasing failures with one cause.

Exploratory testing

Exploratory testing is unscripted testing in which the tester designs and runs checks at the same time, guided by the product, their experience and the last result. It finds the surprises that written test cases miss. It works best time-boxed, with a stated focus and notes kept as the session goes.

Acceptance testing

Acceptance testing decides whether a system is good enough to accept: does it do what the business asked for, in the way that was agreed? It is judged against acceptance criteria rather than code, and usually run by or on behalf of the people who will own the product.

User acceptance testing (UAT)

User acceptance testing, or UAT, is acceptance testing carried out by the people who will actually use the software, to confirm it fits the way they work. Testers follow realistic scenarios rather than scripts, and the outcome is a decision: accept, reject, or accept with conditions written down.

Accessibility testing

Accessibility testing checks that disabled people can use a product, normally against the WCAG success criteria, using a keyboard, a screen reader and an automated rules engine. W3C is clear that tools cannot check everything, so part of every assessment needs human judgement.

Visual regression testing

Visual regression testing compares a screenshot of a page or component against an approved baseline image and reports the differences. It catches broken layouts, wrong spacing and missing styles that functional tests pass straight through. Fonts, animation and changing content cause false alarms unless the comparison is tuned.

Performance testing

Performance testing measures how quickly and efficiently a system responds under a defined workload: page load, response times, throughput and resource use. It answers whether the product is fast enough to use, which is a separate question from whether it is correct.

Load testing

Load testing is performance testing at an expected level of concurrent use, to see whether response times and error rates hold. Stress testing pushes past that level to find where the system breaks, and soak testing keeps a normal load running for hours to expose leaks and slow degradation.

Security testing

Security testing looks for weaknesses an attacker could use, such as broken access control, injection, weak authentication or exposed data. Most teams start from the OWASP Top 10 and work through each risk against their own product, because generic scanners miss anything that depends on business rules.

Usability testing

Usability testing evaluates how effectively, efficiently and comfortably real users can finish tasks with a product. It is done by watching people, not by checking rules, and it answers a different question from accessibility testing: not whether the product can be used, but whether it is easy to use.

Cross-browser testing

Cross-browser testing checks that a site behaves the same across browser engines, chiefly Chromium, Gecko in Firefox and WebKit in Safari, and across screen sizes. Rendering, form controls and newer CSS features still differ between engines, so a page that is perfect in one can be broken in another.

API testing

API testing checks a service's endpoints directly, with no user interface involved: status codes, response bodies, error handling, authentication and limits. Because it skips the browser it is fast and precise, which makes it the natural place to check business rules and changes to a contract between services.

Agentic and AI testing terms

These words come from the newer wave of AI testing tools. Vendors use them loosely, so it helps to be precise about what each one claims.

AI agent

An AI agent is a program that pursues a goal by choosing its own steps, using tools such as a browser, a shell or an API, and reacting to what it observes. In testing, that means deciding what to check, running the checks and interpreting results, rather than replaying fixed instructions.

Agentic testing

Agentic testing is software testing carried out by AI agents that plan, run and judge checks themselves, instead of replaying scripts a person wrote. The agent reads the change and the requirements, works out what is at risk, drives the real product and reports what it found.

Autonomous testing

Autonomous testing is testing that runs without anyone writing or maintaining test cases: the system works out what to check from the product, the requirements or real usage. The word describes how independent a tool is, not how it works, and tools differ enormously in practice.

AI-assisted testing

AI-assisted testing keeps a person in charge and uses a model for the tedious parts: drafting test cases from a ticket, generating selectors, explaining a failure or repairing a broken script. The tests still exist as code or written steps that someone reviews, owns and can run again unchanged.

Browser agent

A browser agent is an AI agent that drives a real browser, clicking, typing, scrolling and reading the rendered page, instead of calling an API. It is how an agent meets the product the way a user does, including layout, focus order and anything that only exists after JavaScript runs.

Self-healing tests

Self-healing tests repair themselves when the interface changes: if a selector no longer matches, the tool finds the element by other attributes and updates the test. That cuts maintenance after cosmetic changes, but it can also hide a real regression by quietly clicking something that was never meant to move.

AI test generation

AI test generation uses a model to write test cases or test code from an input such as a requirement, a design, a recorded session or the running application. The output always needs review, because a generated test can be confidently wrong about what the product was supposed to do.

Large language model (LLM)

A large language model is a model trained on text to predict likely continuations, which is why it can read requirements, summarise a change and write code or test steps. It is strong at language and pattern-matching, and it has no way of knowing whether a claim about your product is true.

LLM-as-judge

LLM-as-judge is the practice of asking a language model to score or grade an output rather than comparing it with an expected value. It is useful where correctness is fuzzy, such as tone or summary quality, but its verdicts drift between runs, so it should never be the only basis for a release decision.

Hallucination

A hallucination is model output that is fluent, plausible and wrong: an invented API, a citation that does not exist, a test report describing a check that never ran. In testing it is dangerous in one direction especially, because a false pass is far harder to notice than a false failure.

Non-determinism

Non-determinism is when the same input does not always produce the same output. Language models are non-deterministic by design, so an AI-driven test can pass one run and fail the next. It is the main reason AI results need a deterministic check, or a second independent run, before anyone relies on them.

Evaluation (eval)

An evaluation, or eval, is a repeatable test suite for an AI feature: a fixed set of inputs, a way of scoring the outputs, and a threshold you expect to hold. Evals are how a team tells whether a prompt, model or retrieval change made the product better or worse.

Prompt injection

Prompt injection is an attack in which content a model reads, such as a web page, a support ticket or a file, carries instructions that change what the model does. OWASP ranks it the top risk for LLM applications, and any agent that browses real content needs limits that hold when the content lies.

Human in the loop

Human in the loop means a person reviews or approves what an automated system produces before it counts. In QA it marks the boundary between machine work and accountability: agents can gather evidence and raise issues, but a named person still decides whether the release goes out.

Knowing the definition is one thing — seeing this kind of evidence on your own release is another, so join the private pilot waitlist to find out firsthand.

Work email only. We keep your email, team size, plan choice and the page you joined from, only to contact you about the ShipperAG pilot. No spam.

Practice and process

The vocabulary of getting a release out of the door: what a team agrees in advance, what it measures, and what it keeps afterwards.

Acceptance criteria

Acceptance criteria are the conditions a piece of work must satisfy before stakeholders accept it, and they are written before the work starts. Useful criteria are specific enough to test: what a user can do, what the system must refuse, and what happens at the edges.

Definition of done

A definition of done is the team's shared standard for when any piece of work is finished, covering things like tests written, review passed, documentation updated and accessibility checked. It applies to every item, where acceptance criteria describe only one.

Test plan

A test plan is the document that says what will be tested, how, by whom, when, and what would make the team stop. On a small team it can be a paragraph in the ticket; what matters is that scope and stopping conditions exist in writing before testing begins.

Test case

A test case is a single check written as preconditions, inputs, actions, expected results and postconditions. The expected result is the important half: without one you have a list of steps and no way to say whether what happened was right or wrong.

Test coverage

Test coverage is the proportion of some defined set, such as lines of code, branches, requirements or user journeys, that your tests actually exercise. High code coverage tells you the code ran while tests were running; it does not tell you that anyone checked the result was correct.

Flaky test

A flaky test passes and fails on the same code, usually because of timing, test order, shared data or a dependency outside the test. Flaky tests cost more than they look: a team learns to re-run them, and then misses the day the failure was real.

Shift left

Shift left means moving testing earlier in the delivery cycle, so defects are found while a change is being designed and written rather than in a release window. In practice it looks like testable requirements, checks that run on every pull request, and a developer who can run them locally.

Quality gate

A quality gate is a rule a change must satisfy before it moves to the next stage: tests green, coverage not falling, no critical vulnerabilities, no new accessibility violations. A gate only works if it genuinely blocks, and if failing it is rare enough that nobody overrides it out of habit.

Continuous integration (CI)

Continuous integration is the practice of merging everyone's changes into the shared codebase at least daily, with an automated build and test run on each merge. It keeps integration problems small, and gives testing a dependable trigger: every change is checked, not just the memorable ones.

Staging environment

A staging environment is a deployed copy of the product that matches production as closely as possible, used for testing before release. The differences that look harmless, such as test data, a smaller database or missing third-party keys, are exactly where staging stops predicting what production will do.

Preview environment

A preview environment is a temporary deployment created automatically for one branch or pull request, with its own URL. It lets reviewers and testers use the change before it is merged, which is what makes testing every change, rather than every release, practical for a small team.

Canary release

A canary release rolls a new version out to a small share of users first, watches what happens, and only then sends everyone to it. It limits the damage a bad release can do, and gives real production signal that no test environment can reproduce.

Rollback

A rollback returns a system to the previous working version after a bad release. It is only as good as the parts you cannot undo: a database migration, an email already sent, a payment already taken. Knowing which changes are reversible belongs in the release plan, not the incident.

Release readiness

Release readiness is the judgement that a change is safe to ship: it does what was asked, known risks are covered or consciously accepted, and someone has seen the evidence. It is a decision rather than a test result, so it needs a written standard to be the same every time.

Sign-off

Sign-off is a named person recording that they accept a release on behalf of the business, with a date, a scope and a note of what they relied on. It is the difference between "the tests passed" and "we decided to ship", and it is what stakeholders ask for afterwards.

Evidence

Evidence in QA is the record of what was actually checked and what happened: screenshots, request and response logs, steps to reproduce, timestamps and the build it ran against. Without it a result is only a claim, and months later nobody can tell which claim covered which release.

Traceability

Traceability is the ability to follow the links between related work: this requirement to that test, that test to this defect, that defect to the release that fixed it. It is what lets a team answer whether every requirement was checked, and what a proposed change would affect.

Test automation

Test automation is the use of software to run tests and compare results, instead of a person doing it by hand. It is fast and repeatable for the paths somebody has already written down, and the cost sits in maintenance, because every interface change can break tests that were working.

Standards, specifications and tools

The specifications and tools that come up most often in QA conversations, each with its primary source.

WCAG 2.2

WCAG 2.2 is the W3C's Web Content Accessibility Guidelines, a Recommendation published in December 2024 that sets testable success criteria under four principles: perceivable, operable, understandable and robust. Conformance is claimed at level A, AA or AAA, and most laws and contracts ask for AA.

EN 301 549

EN 301 549 is the European accessibility standard for ICT products and services, used as the reference for public sector rules and the European Accessibility Act. Version 4.1.1, published in September 2026, aligns with WCAG 2.2 and adds requirements WCAG does not cover.

OWASP Top 10

The OWASP Top 10 is a consensus list of the most critical security risks to web applications, published as an awareness standard by the Open Worldwide Application Security Project. In the 2025 edition, broken access control is first, and it is a class of flaw that functional tests rarely catch.

OWASP Top 10 for LLM Applications

The OWASP Top 10 for LLM Applications is the equivalent list for products built on language models, covering prompt injection, sensitive information disclosure, improper output handling and excessive agency among others. It is the usual starting point for deciding what to test in an AI feature.

axe-core

axe-core is an open-source accessibility rules engine, maintained by Deque Systems, that runs inside a page and reports violations of WCAG and best-practice rules. It is the engine behind many browser extensions and test integrations, and it lists uncertain cases separately for a human to review. See the 32 WCAG criteria it has no rule for.

Playwright

Playwright is an open-source browser automation and testing framework from Microsoft that drives Chromium, Firefox and WebKit through one API, in TypeScript, Python, .NET or Java. It is the layer many AI testing tools use to reach a real browser, rather than a testing approach itself.

Selenium

Selenium is the long-standing open-source project for automating browsers, made up of the WebDriver language bindings, the record-and-playback Selenium IDE, and Selenium Grid for running tests across many machines and browser versions. Its WebDriver component implements the W3C protocol of the same name.

WebDriver

WebDriver is the W3C standard remote control interface for browsers: a platform-neutral protocol that lets an external program find elements, click, type and read state. It became a W3C Recommendation in 2018, which is why one test suite can drive several different browsers.

Headless browser

A headless browser is a real browser engine running with no visible window, driven entirely by code. It renders pages and executes JavaScript as usual, which makes it the standard way to run browser tests on a server, though a few behaviours, such as some GPU rendering, can still differ.

Device emulation

Device emulation is a desktop browser pretending to be a phone or tablet, by simulating the viewport size, pixel ratio, user agent and touch input. It is quick and catches most layout and touch-target problems, but it is not a real device and cannot prove how an actual phone browser behaves.

Model Context Protocol (MCP)

The Model Context Protocol is an open standard for connecting AI applications to outside tools and data through one common interface, so any compatible client can use any compatible server. In testing, it is how an agent is handed the browser, repository or tracker it needs.

Sources, each opened while writing this page: W3C, WCAG 2.2, WebDriver and WAI, Selecting Web Accessibility Evaluation Tools; ETSI, EN 301 549 V4.1.1; OWASP Top 10:2025 and OWASP Top 10 for LLM Applications; Deque, axe-core; Playwright; Selenium; Model Context Protocol; the ISTQB Glossary for testing terms; Martin Fowler, Continuous Integration; Danilo Sato, Canary Release. Checked 20 September 2026.

FAQ

The questions a glossary does not answer

What is the difference between QA and testing?

Testing is one activity inside QA. Testing runs the product and checks what it does. Quality assurance is the wider practice of deciding what quality means for this product, agreeing acceptance criteria, choosing what is worth checking, and making sure the release decision is evidence-based. A team can test a great deal and still have no quality assurance, if nobody decided what correct means.

What is the difference between automated testing and autonomous testing?

Automated testing runs tests a person wrote and maintains: the same steps every time, and somebody has to update them when the product changes. Autonomous testing works out what to check from the requirements and the change, runs those checks, and reports what it could not cover. The difference is not who presses the button, it is who decides what to test.

What is the difference between verification and validation?

Verification asks whether the product was built to the specification. Validation asks whether the specification described what people actually needed. A release can pass verification completely and still fail validation, which is the case where every test is green and the feature is wrong.

What is a flaky test?

A flaky test is one that passes and fails on the same code without anything changing, usually because of timing, shared state, network conditions or test order. Flaky tests are worse than missing tests: a team learns to re-run them until they go green, and then does the same to a real failure.

What does agentic mean in AI testing?

Agentic means the AI decides the steps rather than following them. An agentic testing tool reads the context, plans which checks a change needs, drives the product itself, reacts to what it finds and reports the result. An AI-assisted tool, by contrast, helps a person write or repair tests that the person still runs.

Free 45-day trial · waitlist open

From definitions to evidence

ShipperAG checks a release against what it is supposed to do and hands your team evidence it can sign: see how it works. It is in a private pilot, and joining the waitlist is how a team becomes a product partner.

  • Free 45-day trial, no card
  • Up to 20 release checks on your product
  • Direct line to the founders

Work email only. We keep your email, team size, plan choice and the page you joined from, only to contact you about the ShipperAG pilot. No spam.