Blog · AI and quality

AI writes the code. Who checks it?

Most developers now use AI coding tools, yet more of them distrust the output than trust it. The gap between writing code faster and knowing it works is where releases go wrong. This piece looks at that gap, how teams can close it, and the thinking behind ShipperAG.

Published · ShipperAG team · Free to republish (CC BY 4.0)

The trust gap, in numbers

The Stack Overflow Developer Survey 2025 found that 84% of respondents use or plan to use AI tools in their development process. In the same survey, 46% said they actively distrust the accuracy of AI tools, against 33% who trust it, and only 3% said they highly trust the output.

So the tools are everywhere, and the people using them are not sure the results are right. That is a healthy instinct. It is also a lot of doubt to carry into every release.

Why “almost right” is the expensive kind of wrong

In the same survey, 66% of developers named “AI solutions that are almost right, but not quite” as their biggest frustration. Code like that reads well in review and passes the happy path. It fails at the edges: a boundary value, an error state, a permission check, a screen reader, a small phone.

There is a quieter problem too. When the same assistant writes the code and the tests, the tests share the code's assumptions. Everything is green, and the one thing nobody wrote down is still wrong.

Check against intent, not against the code

The code is not its own specification. The useful question is not “does this code do what this code does” but “does the product do what we said it should”. That means checking behaviour against the brief, the acceptance criteria, the design and the business rules.

If the spec says free shipping starts at £50, the check is whether a basket of £49.99 pays for shipping and a basket of £50 does not, whoever or whatever wrote the code. Our guide on testing AI-generated code walks through this step by step.

Where AI agents help, and where they must not be trusted

AI agents are good at broad, repetitive checking: walking journeys, trying edge cases and comparing screens on every release, without anyone writing scripts. That is where agentic AI testing earns its place, including testing in a real browser.

But an agent's opinion is not proof. A pass has to come from a check that actually ran. A failure has to reproduce in a fresh session before anyone is paged about it. And a person still decides what ships.

What we are building

ShipperAG applies those rules to every release: AI QA specialists check it against what it is supposed to do, re-check what they find and hand the team evidence it can sign. We are keeping the details close while the private pilot runs. The best way to see it is on your own release.

Free to republish. This article is licensed under Creative Commons Attribution 4.0 (CC BY 4.0). Copy it, share it or adapt it, as long as you credit ShipperAG and link to this page.

Sources: Stack Overflow Developer Survey 2025 (AI section). Checked 20 September 2026.

Free 45-day trial · waitlist open

Close the trust gap

ShipperAG is in a private pilot. Join the waitlist to become a product partner and see it check a real release.

  • Free 45-day trial, no card
  • Up to 20 release checks on your product
  • Direct line to the founders

Work email only. We keep your email, team size, plan choice and the page you joined from, only to contact you about the ShipperAG pilot. No spam.