Short answer. Agentic testing with Playwright means an AI agent drives Playwright itself: Playwright's own Test Agents (planner, generator and healer, from version 1.56) plan, write and repair tests, while Playwright MCP and the Playwright CLI let any coding agent control a browser directly, step by step. Test Agents keep everything as ordinary Playwright code you review and run in CI; MCP suits longer exploratory sessions; the CLI is the lighter option for a coding agent mid-task. ShipperAG is a private pilot where tech and product teams check their own product's releases in real Chromium browser sessions; Firefox, Safari and real devices are not covered yet.
What is agentic testing with Playwright?
Agentic testing with Playwright is testing where an AI agent, not a person, decides the browser steps and Playwright carries them out. The agent works towards a goal, such as "check that a new user can sign up and reach the dashboard", and chooses each action from what is on the page.
Playwright is a good fit because it already does the hard part of browser automation well: it starts Chromium, Firefox or WebKit, waits for pages to be ready, and clicks, types and reads the page reliably. An agent adds judgement on top. It can explore an unfamiliar app, write a test for a flow nobody has scripted yet, or notice that a test failed because a button was renamed rather than because the feature broke.
There are three main ways to connect an agent to Playwright today, all maintained by the Playwright team at Microsoft: Playwright Test Agents, Playwright MCP and the Playwright CLI with skills. They solve different problems.
Playwright Test Agents: planner, generator and healer
Playwright Test Agents are three agent definitions that come with Playwright itself, introduced in version 1.56. Each one does one job in the life of a test.
- Planner. Explores the app and writes a test plan in Markdown, saved in a
specs/folder. People can read and edit the plan before any code exists. - Generator. Turns the Markdown plan into Playwright test files, checking selectors against the live page as it goes.
- Healer. Runs the test suite and repairs tests that fail, for example when a locator no longer matches.
You add them to a project with npx playwright init-agents and a --loop option for your AI client; Playwright's documentation lists VS Code, Claude Code, Codex and opencode. A seed test gives the agents a ready page to start from. The documentation also notes that the agent definitions should be regenerated when you update Playwright, so they pick up new tools.
The result is ordinary Playwright test code that lives in your repository and runs in CI like any other test. That is the main appeal: the agent does the tedious writing and repair, and your team keeps code it can review.
Playwright MCP and the Playwright CLI
Playwright MCP and the Playwright CLI let any AI agent control a browser directly, step by step, rather than only producing test files.
Playwright MCP is a Model Context Protocol server that gives a model browser tools. It works from Playwright's accessibility tree, the same names and roles a screen reader uses, so the model does not need to read screenshots. Its README describes this as faster and less ambiguous than pixel-based input. It can run Chrome, Firefox, WebKit or Edge.
The Playwright CLI gives coding agents short commands such as open, click, fill, snapshot and screenshot, and playwright-cli install --skills adds skills that teach an agent how to use them. This is what people mean by "agentic testing with Playwright CLI skill".
Playwright's own guidance on choosing between them is practical: the CLI is more token-efficient for coding agents because it avoids loading large tool schemas and accessibility trees into the model's context, while MCP suits longer agent loops that need persistent browser state, such as exploratory testing, self-healing tests or long-running autonomous workflows.
Test agents, MCP, CLI and plain scripts compared
Each option puts the agent in a different place. Pick by what you want to own at the end.
| Option | What the agent does | What you keep | Good for |
|---|---|---|---|
| Plain Playwright scripts | Nothing; a person writes every step | Test code | Critical flows that must behave the same on every run |
| Playwright Test Agents | Plans, writes and repairs tests | A Markdown plan and test code | Growing and maintaining a suite with less hand-writing |
| Playwright CLI with skills | Drives the browser with short commands from a coding agent | Whatever the agent writes, often tests | Checking a change while coding, at low token cost |
| Playwright MCP | Drives the browser through tools in a long session | The session's results and any tests it writes | Exploratory checks and longer autonomous runs |
For the wider market of tools built on these ideas, see AI testing tools in 2026 and our guide to AI browser testing.
A practical workflow for teams
Most teams get the best results by letting agents do the writing and exploring, while people decide what "correct" means and review what goes into the suite.
- Write down what should happen. Acceptance criteria, business rules and the main user journeys. An agent without them can only check that pages load and buttons respond.
- Test a preview or staging build. Use test accounts and test data, never production. An agent that can click can also delete, send and pay.
- Let the planner propose, and review the plan. Editing a Markdown plan is much quicker than reviewing generated code.
- Generate tests for the flows that matter most, and review them like any other pull request before they join the suite.
- Use an agent for exploration around the suite: new features, edge cases and pages nobody has scripted.
- Review every healed test. A repaired locator is fine; a repair that changes what the test checks is not.
- Keep evidence. Traces, screenshots and steps for every result, so anyone can see what was really checked.
This split mirrors the one in agentic QA vs test automation: scripts for the handful of flows that must never break, agents for breadth.
Want to see agentic testing with Playwright run against your own product's next release? Start with a free 45-day trial, no card.
Work email only. We keep your email, team size, plan choice and the page you joined from, only to contact you about the ShipperAG pilot. No spam.
Risks to watch
Agents add flexibility, and with it some new ways to get a wrong answer. These are the ones that matter most with Playwright.
- Healing that hides bugs. If a healer "fixes" a test by changing what it asserts, a real regression can turn green. Review healed tests, and treat changed assertions as a red flag.
- Tests that check the wrong thing. Generated tests often confirm what the app does today, not what it is supposed to do. Tie them to written acceptance criteria.
- Passes from the model's word. A pass should come from an observable check, such as page state, a URL or a network response, never from the agent's summary.
- One engine only. Playwright can run Chromium, Firefox and WebKit, but many agent set-ups use only Chromium. A pass there says nothing certain about Safari.
- Cost and speed. Model calls cost money and take time. Keep fast, deterministic scripts for checks that run on every commit.
Where ShipperAG fits
ShipperAG is agentic QA for release sign-off rather than a tool for writing Playwright tests. Its specialists run their checks in real Chromium browser sessions, against a preview or staging build you connect, and check it against your requirements, design system and business rules. If you are weighing that up against keeping your own suite, read agentic QA vs Playwright.
- Beyond the browser. Alongside browser checks, specialists cover APIs, data integrity, security and more of the 13 quality areas.
- Checks decide, not the model. Language models help read context and plan; deterministic checks decide pass or fail.
- Every finding re-run. An independent verifier repeats each finding in a fresh session before it reaches you.
- Evidence you can sign. Results are sealed in a hash-chained evidence ledger for release sign-off.
What ShipperAG does not cover today: Firefox and Safari (WebKit) rendering, and real devices. Phone and tablet sizes are checked with Chromium's device emulation. When a release needs more, the report lists it under "Not covered".
It works alongside your own Playwright suite rather than replacing it. See how ShipperAG works and the AI QA specialists.
Sources: Playwright: Test Agents; Playwright release notes (version 1.56); Playwright MCP (Microsoft, GitHub); Playwright CLI (Microsoft, GitHub); Playwright: Browsers. Checked 2 October 2026.