Blog · From the team

Why we're building ShipperAG, and what comes next

Tech and product teams ship faster than ever, and more of their code is written with AI. What has not kept pace is the answer to one question before every release: how do we know it works the way it should? ShipperAG is our answer. This note explains why we started, what we believe and what our private pilot will explore.

Published · ShipperAG team · Free to republish (CC BY 4.0)

The question every release has to answer

“Is this release ready?” Every team answers that question, usually under time pressure. The answer is often a mix of green checks in CI, a quick click through staging and a message in a team channel that says “looks good to me”.

That works until it does not. A pricing rule breaks, a checkout step fails on a phone, a form stops working for keyboard users. Nobody decided to skip those checks. There was simply no time to run them on every release.

Tests prove that code does what its tests say. They rarely prove that the product does what it is supposed to do. That gap is where we started.

AI changed how code is written, not how it is checked

According to the Stack Overflow Developer Survey 2025, 84% of developers use or plan to use AI tools in their development process, yet more of them distrust the accuracy of those tools (46%) than trust it (33%). The biggest single frustration, cited by 66%, is AI solutions that are “almost right, but not quite”.

“Almost right” is exactly the kind of problem that survives a quick look and then fails for a real customer. More code, written faster, needs more checking, not less. We wrote a practical guide on how to test AI-generated code for teams dealing with this today.

What we believe

  • Correct means what you asked for. A release should be checked against the requirements, the design and the business rules, not against the code itself.
  • Evidence, not opinions. An AI model saying “this looks fine” is not proof. A result should come from a check that actually ran, and a problem should be reproduced before anyone is asked to act on it.
  • Honest about gaps. A useful report says what was checked, what failed, what needs a decision and what was not covered.
  • People make the call. Agents can do the legwork. A person decides whether a release ships.

What we are building

ShipperAG is a team of AI QA specialists that checks each release of a website, app or product against what it is supposed to do, re-checks what it finds and hands the team evidence it can sign. You can read the idea behind agentic AI testing in our guide.

We are deliberately not sharing much more yet. The product is changing quickly as we test it on real releases, and we would rather show it working than describe it.

What the private pilot will explore

The pilot is for a small number of tech and product teams at SaaS companies, start-ups and e-commerce businesses. Together we will work on real releases and a few open questions:

  • Which checks do teams trust enough to rely on before a release?
  • What evidence do product owners and stakeholders actually read?
  • How should an agentic QA tool ask for help when requirements are unclear or contradict each other?

If those questions matter to your team, join the waitlist. We will share more here as the pilot progresses.

Free to republish. This article is licensed under Creative Commons Attribution 4.0 (CC BY 4.0). Copy it, share it or adapt it, as long as you credit ShipperAG and link to this page.

Sources: Stack Overflow Developer Survey 2025 (AI section). Checked 20 September 2026.

Free 45-day trial · waitlist open

Help shape what comes next

We are opening the private pilot to a small number of tech and product teams. Join the waitlist to become a product partner and try it on a real release.

  • Free 45-day trial, no card
  • Up to 20 release checks on your product
  • Direct line to the founders

Work email only. We keep your email, team size, plan choice and the page you joined from, only to contact you about the ShipperAG pilot. No spam.