Regression testing on every pull request is the chore ContextQA now wants an AI to take off engineers’ hands. On 5 October 2026 the Austin, Texas-based test automation company launched Ship, which it calls an “Autonomous Quality Engineer”. Ship reads each pull request, works out which user flows the change could break, runs the relevant regression tests against the deployed app before merge, and keeps working after release by reproducing bug reports and mining production sessions for new tests.
The launch lands at a moment when AI coding tools are producing more changes than human reviewers and QA teams can comfortably check. It also arrives with unusually open pricing: one credit worth $0.025, a $49 Starter plan and a published rate for every workflow. That makes it possible to work out what automated regression testing on pull requests would actually cost a real team, rather than guessing from a “contact sales” button.
This article explains what Ship does, how its pull request workflow compares with older regression testing techniques and with AI code reviewers, what the credit maths looks like for three kinds of team, which of its headline numbers are illustrative, and how to run a fair 15-day trial.
Table of contents
- What ContextQA Launched for Regression Testing on 5 October
- How Ship Runs Regression Testing on Pull Requests
- Beyond the Pull Request: Regression Testing From Bugs and Production
- Why Regression Testing Is Under Pressure From AI-Written Code
- Smarter Regression Testing Is Not New: Where Ship Fits
- Ship Pricing: What Regression Testing Costs in Credits
- Ship vs AI Code Review and Test-Selection Tools
- The Numbers Ship Advertises, and How to Read Them
- Risks and Open Questions for Automated Regression Testing
- How to Trial Ship’s Regression Testing in 15 Days
- Who Should Consider Ship for Regression Testing
- Regression Testing With Ship: FAQs
- References
What ContextQA Launched for Regression Testing on 5 October
Ship is a new, self-serve regression testing product from a company that already sells an enterprise test automation platform. Here is what was announced, by whom, and what the label “autonomous quality engineer” is meant to cover.
The launch post and the demo
ContextQA announced Ship in a LinkedIn article by Swati Uniyal, published at 13:33 UTC on 5 October under the line “Your Autonomous Quality Engineer. Always testing.” Two hours later the AI news account TestingCatalog posted a demo clip with the one-line summary that most coverage has since repeated: “Ship can run on pull requests by testing user flows affected by changes in the deployed app and finding regressions before release.”
The Ship product page opens with a scrolling terminal feed labelled “demo stream”. It shows Ship detecting a new deployment, running 148 tests across checkout, auth and billing, finding two defects in PR #1234, creating two regression test cases for that PR, reproducing a checkout bug that appears after a session expires, and tracing the root cause to a stale token on line 87 of auth/refresh.ts.
Each of the three workflow panels underneath the feed is marked “illustrative workflow”. That label matters when reading the rest of the page, and we come back to it below.
Who ContextQA is
ContextQA is not new to test automation or regression testing. Its main platform sells AI-generated test cases, self-healing tests that repair broken selectors when an interface changes, parallel runs across browsers and devices, and testing of AI agents‘ responses, guardrails and tool calls. It also exposes about 50 testing tools over MCP, so developers can run suites from Cursor, Claude Code or Codex.
An IBM case study describes the company as based in Austin, Texas, quotes founder and CEO Deep Barot, and reports that ContextQA used IBM watsonx.ai to migrate and automate 5,000 existing test cases within minutes.
Ship lives on its own subdomain with its own trial and pricing, but it shares the parent platform’s currency. The FAQ confirms that Ship credits cost the same $0.025 as ContextQA’s existing AAT credits.
What “autonomous quality engineer” means
The phrase does a lot of work, and the FAQ answers the obvious question directly. “Is Ship another test automation tool? No. Ship is an autonomous quality engineer. It decides what needs attention, runs the appropriate quality workflow, investigates failures, and delivers actionable results without waiting for someone to manually create and execute every test.”
In practice that means three event-driven workflows. A pull request, a bug report or a production anomaly triggers the regression testing work, rather than a test plan that a person schedules. The rest of this article takes them in turn, starting with the one TestingCatalog highlighted.
How Ship Runs Regression Testing on Pull Requests
The pull request workflow is the one most teams will evaluate first. Ship’s page sums it up in one sentence: it “analyzes each PR, understands its blast radius, identifies what needs testing, and runs the right regressions before merge.” Each step is worth unpacking, because each one is a place where regression testing can go right or wrong.
Mapping the blast radius
Before Ship can choose what regression testing to run, it needs a model of the application. A one-off code analysis reads the whole repository and “maps how it fits together and generates tests from it”. A crawl explores the running application to map its pages and flows. Together they give Ship a picture of which journeys depend on which code.
ContextQA’s older Impact Analysis product, which Ship’s PR workflow closely resembles, describes the same idea as a live knowledge graph. Requirements, test cases and the application sit in one model, so a change can be traced “to every test and journey it can reach, with no manual tagging”.
In the illustrative example on the Ship page, a three-file pull request called fix/checkout-session, adding 42 lines and removing 18, is mapped to three affected flows: checkout, payments and auth.
Targeted regression testing before merge
Once the affected flows are known, Ship runs only the regression testing those flows need, rather than the whole suite. The same example shows 12 of 12 targeted regressions completed and one regression caught before merge, “Expired sessions block the payment step”, flagged as P1.
The Impact Analysis page gives a comparable picture with different numbers. It runs 25 of 340 tests for one pull request, skips the 315 that were not affected, and posts the result back to GitHub, GitLab or Jenkins as a status check, so that “a high-risk change cannot merge unreviewed”.
Testing the deployed app, not the diff
The design choice that sets Ship apart from most PR tooling is that it tests the deployed product. A test run, in the pricing page’s words, “executes your test in a real browser, with video, trace and screenshots”. That is closer to what a human tester does on a staging environment than to what a code reviewer does on a diff.
It is also the basis of ContextQA’s argument that Ship is a different category from AI code review. The trade-off is practical: browser-based regression testing needs a deployed environment for every change it checks. Confirm you have preview or staging environments per branch before planning a trial.
New tests written from pull requests
Ship also writes tests while it analyses pull requests, and the pricing page lists “tests generated from PRs” as free. Over time that should make regression testing coverage grow with the codebase instead of drifting behind it, provided someone reviews what gets added.
The table below summarises the three workflows, what triggers each one, and what each costs.
| Workflow | Trigger | What Ship does | What you get | Credit cost |
|---|---|---|---|---|
| Change testing | A pull request or new deployment | Maps the blast radius, then picks and runs targeted tests in a real browser | Results per affected flow, defects caught before merge, new tests | 10 per PR analysis, plus 1 per 4 test steps |
| Bug investigation | A ticket in Linear, Jira, Slack or a support queue | Reproduces the bug, records a replay and network trace, looks for the cause | Reproduction steps, evidence, likely cause, suggested fix, a new test | 15 per bug |
| Production signals | Anomalies in real user sessions | Reproduces suspicious behaviour and turns it into a test | A permanent test that runs on every future PR | No separate rate published |
Beyond the Pull Request: Regression Testing From Bugs and Production
Pull requests are only one source of regression testing work. Ship’s other two workflows start from things that have already gone wrong, and both feed back into the regression testing suite.
Bug reproduction from the tracker
The second workflow starts when a bug lands in Linear, Jira, Slack or a service desk. Ship picks it up, reproduces it, captures “step-by-step evidence” such as a session replay and network trace, and investigates the likely root cause.
In the illustrative example, a customer report that “the spinner keeps going after my session expires” is reproduced in 38 seconds and traced to a null token on line 87 of auth/refresh.ts. When Ship confirms a bug, it adds a test for it, so the same failure is checked on future changes. That is the classic purpose of regression testing: a bug fixed once should stay fixed.
Production sessions feed regression testing
The third workflow is the most ambitious. Ship “analyzes real user sessions for anomalies, reproduces suspicious behavior and turns what it learns into permanent regression coverage”.
The example shows 1,284 sessions analysed in 15 minutes, one anomaly (repeated payment attempts), and a new test called “Checkout recovers expired sessions” added to the suite to run on every future PR. This is where regression testing stops being a fixed list and starts learning from production. It is also the part of the product that raises the most questions about data access, which we cover in the risks section.
Hand-offs to Claude and Codex
Ship suggests fixes but does not try to replace a coding agent. Its bug panel carries two buttons, “Send to Codex” and “Send to Claude”, and the FAQ frames the relationship plainly: “Ship isn’t trying to replace coding agents. It gives them better inputs.”
For teams already using AI coding assistants, the useful promise is a hand-off that contains a reproduction, a trace and a likely cause. A coding agent given a vague ticket has to establish what happened first, which is slow and error-prone.
Why Regression Testing Is Under Pressure From AI-Written Code
ContextQA’s launch post makes its case with outside research, and the figures hold up when checked against the original sources. Together they describe a widening gap between how fast code is written and how carefully it is checked.
The verification gap
Sonar’s State of Code Developer Survey, released on 8 January 2026 and based on more than 1,100 developers, found that 96% do not fully trust that AI-generated code is functionally correct. Yet only 48% say they always check AI-assisted code before committing it.
The same survey put AI’s share of committed code at 42%, a figure developers expect to reach 65% by 2027. And 38% said reviewing AI-generated code takes more effort than reviewing code written by a colleague.
Throughput up, stability down
Google’s 2025 DORA report, based on nearly 5,000 technology professionals, found that 90% use AI at work and more than 80% believe it has made them more productive. It also found that AI adoption now has a positive link with delivery throughput, but “does continue to have a negative relationship with software delivery stability”.
DORA’s explanation reads almost like a brief for automated regression testing. Without “strong automated testing, mature version control practices, and fast feedback loops, an increase in change volume leads to instability.”
Here are the four figures that frame the problem, from the two surveys above:
More changes, the same QA team
The launch post also cites an April 2026 Gartner journey guide that calls for a “continuous quality” strategy. The common thread is volume. When AI multiplies the number of pull requests, regression testing has to scale with changes rather than with headcount, and running the full suite on every change becomes either slow or expensive.
We have seen the same pressure from another angle. Google paused part of its open source bug bounty after a flood of AI-generated reports, because human reviewers could not keep up with machine-made volume.
Smarter Regression Testing Is Not New: Where Ship Fits
Running only the tests a change can affect, often called selective regression testing, is a well-established idea with a decade of engineering behind it. What is new in Ship is the unit it selects on, and how much else it bundles around the selection.
Meta’s predictive test selection
In 2018 Meta, then Facebook, described a system that learned which tests to run for each code change. It reported catching “more than 99.9 percent of all regressions before they are visible to other engineers in the trunk code” while “running just a third of all tests that transitively depend on modified code”.
The result, Meta said, was to “double the efficiency” of its testing infrastructure. That is the core economic argument for selective regression testing: most tests, for most changes, tell you nothing new.
Test selection products for regression testing today
Several products already sell selective regression testing to everyone else. Develocity Predictive Test Selection “identifies the tests relevant to a code change and runs only those tests” for Gradle and Maven builds.
Datadog’s Test Impact Analysis, formerly Intelligent Test Runner, “automatically selects and runs only the relevant tests for a given commit” across .NET, Java, JavaScript, Swift, Python, Ruby and Go. And CloudBees bought Launchable, the test-selection start-up co-founded by Jenkins creator Kohsuke Kawaguchi, in August 2024.
Flows, not files
Those tools mostly choose from tests you already have, at the level of test classes and code dependencies. Ship works at the level of user journeys in the deployed app, writes missing tests itself, and adds bug reproduction and production analysis on top.
That is a bigger claim and a harder one to verify. File-level test selection can be checked against a full run of the suite. Flow-level regression testing depends on how closely Ship’s model of the application matches reality, which only a trial on your own code can show.
Ship Pricing: What Regression Testing Costs in Credits
Unusually for a launch-day product, Ship publishes complete pricing. Everything runs on one credit worth $0.025, and every plan includes unlimited users. The plans differ in monthly credits, repositories and parallel regression testing runs.
| Plan | Price | Credits | GitHub repositories | Parallel runs | Notable extras |
|---|---|---|---|---|---|
| Free trial | $0 for 15 days | 3,000 one-time | 1 | 2 | Every Starter feature |
| Starter | $49 a month | 2,000 a month | 1 | 5 | Extra credits at $0.025 |
| Pro | $249 a month | 12,000 a month | Multiple | 10 | Real mobile devices “coming soon” |
| Enterprise | Custom | By contract | Unlimited | Custom | SSO, on-premise, dedicated support |
What a credit buys
Every plan uses the same rates. The dollar column is our conversion at the list price of $0.025 a credit; Starter’s included credits work out at $0.0245 each and Pro’s at about $0.021.
| Workflow | Credits | Cost at $0.025 |
|---|---|---|
| Test run | 1 per 4 steps, so a 30-step test is 8 | $0.20 for a 30-step test |
| PR impact analysis | 10 per pull request | $0.25 |
| Bug analysis | 15 per bug | $0.375 |
| Code analysis | 1,000 per repository, usually once | $25 |
| Crawl | 500 per crawl | $12.50 |
| Requirement upload | 20 per upload | $0.50 |
| Ask a document | 5 per question | $0.125 |
| Chrome recording | 2 per test case | $0.05 |
| Root cause and autofix, tests from PRs, knowledge sync | Free | $0 |
Plan credits reset every month and do not roll over. Extra credits never expire, and plan credits are always spent first. A run that stops early is charged only for the steps it executed, and a failed test costs no more than a passing one, with root cause analysis included.
Three monthly scenarios
The scenarios below use only the published rates, but the assumptions are ours. We assume each analysed pull request triggers 12 targeted tests of 30 steps each, so one PR costs 10 + (12 × 8) = 106 credits, or $2.65 at list price. Each scenario also includes 10 bug analyses at 15 credits each.
- Promotion PRs only. 20 promotion PRs a month (2,120 credits) plus 10 bugs (150) is 2,270 credits. On Starter that is $49 plus 270 extra credits ($6.75): $55.75 a month.
- ContextQA’s own estimate. The pricing calculator’s default month is 450 test runs of 30 steps (3,600 credits), 5 PR analyses (50) and 10 bugs (150): 3,800 credits. It recommends Starter at $94 a month, $49 plus 1,800 extra credits.
- Every feature PR. 200 PRs a month (21,200 credits) plus 10 bugs is 21,350 credits. On Pro that is $249 plus 9,350 extra credits ($233.75): $482.75 a month.
The chart shows how far apart those monthly bills are:
Why the FAQ steers you to promotion PRs
The FAQ’s answer to “Which pull requests should Ship analyse?” is revealing: “Any you like, at 10 credits each. Many teams analyse every promotion PR, for example dev to QA, and run the impacted tests there, which keeps credit use predictable as the number of feature PRs grows.”
The analysis itself is cheap at $0.25. The cost scales with the browser tests each analysis triggers. A team that wants regression testing on every feature branch should budget for the test runs, not the analysis, and should check how many tests a typical PR actually selects during the trial.
Ship vs AI Code Review and Test-Selection Tools
ContextQA’s FAQ names two competitors directly, CodeRabbit and Greptile, and draws the line this way: “Code review tools primarily reason about code. Ship tests the deployed product like a quality engineer would.” The table puts Ship next to those reviewers and the test-selection tools described earlier.
| Tool | Category | What it examines | When it runs | Pricing model |
|---|---|---|---|---|
| Ship by ContextQA | Autonomous QA | Behaviour of the deployed app across user flows | PRs, deploys, bug tickets, live sessions | Credits at $0.025, unlimited users, from $49 a month |
| CodeRabbit | AI code review | The code changed in each PR | Every PR | Per developer: $24, $48 or $72 a month, billed annually |
| Greptile | AI code review | The code changed, with codebase context | Every PR | $30 per seat a month with 50 credits per seat |
| Develocity | Test selection | Existing tests in Gradle and Maven builds | Every build | Part of the Develocity platform |
| Datadog | Test selection | Existing tests in seven languages | Every commit | Part of Datadog’s test tooling |
Code review reasons about code
The categories are starting to overlap. CodeRabbit’s $72 Advanced tier advertises “Blast radius and architectural impact analysis”, the same phrase Ship uses, but applied to the code rather than to a running application.
In practice the two are complementary. A reviewer can spot a missing null check on line 87 of a file. Browser-based regression testing shows that the checkout actually stalls when a session expires. Most teams will want both code review and regression testing, at different points in the pipeline.
Seat pricing vs work pricing
The commercial difference is just as clear. A ten-developer team on CodeRabbit Team pays $480 a month on annual billing, and on Greptile Pro it pays $300 a month with 500 review credits.
Ship charges nothing per seat, so its bill tracks how much regression testing you ask it to do rather than how many people log in. That favours teams with a steady, predictable volume of tests, and it can hurt teams that push hundreds of pull requests a month through browser tests.
The Numbers Ship Advertises, and How to Read Them
Ship’s product page carries four headline figures. None comes with a methodology, a sample size or a date, so treat them as vendor claims rather than measured results.
The sample release report
Below the figures sits a sample “Release Quality Report” for Release 284. It shows a quality score of 92 out of 100, 31 code changes analysed, 47 impacted tests identified, 3 regressions caught, 1 production anomaly reproduced, 2 regression tests added and $18.4K of “estimated defect cost avoided”.
It is a useful picture of the reporting Ship intends to produce. Nothing on the page says it describes a real customer release, and the $18.4K figure depends entirely on what a team assumes a defect costs.
What the launch post itself concedes
To its credit, the launch article is careful about this. It calls the Sonar figures “useful industry context, rather than evidence of Ship’s performance”, and says the checkout example “should not be read as a measured customer outcome”.
It also recommends judging Ship on “reproduction effort, the clarity of failure evidence, coverage of changing journeys and the time required to move from a report to an actionable investigation”, against a baseline. That is the right standard for any automated regression testing product, and the trial plan below applies it.
Who the testimonials come from
The four customer quotes on the Ship page come from people at Halight, Clari, Skillibrium and Codexitos. Three of the four also appear on ContextQA’s main site praising the existing platform. They read as endorsements from established customers, not independent reviews of a product launched that day.
Risks and Open Questions for Automated Regression Testing
None of this makes Ship a bad bet for regression testing. It does mean a trial should test the weak points as hard as the demo tests the strong ones.
Flaky tests and false confidence in regression testing
Browser tests are prone to flakiness. Timing, test data and third-party services can make a test fail without any real defect, and an autonomous system that writes tests can add flaky ones as quickly as useful ones. Ask how Ship tells a genuine regression from an intermittent failure, and whether it quarantines unstable tests.
The opposite failure is quieter. A green targeted run only means the flows Ship chose were fine. If its blast-radius map misses a dependency, regression testing will pass a change that breaks something it never looked at. Run the full suite on a schedule as a backstop.
Access to code, environments and sessions
To do its job Ship needs read access to your repositories, a deployed environment for each change, your issue tracker and, for the production workflow, real user sessions. That is a lot of sensitive data flowing to a third party.
The Ship page links to a data processing agreement and a data privacy policy, but SSO and on-premise deployment are Enterprise-only. UK teams handling personal data should check where session data is processed and stored, and how it is masked, before switching on production analysis. Treat the integration like any other cybersecurity review of a supplier with production access.
Credits that expire
Monthly plan credits do not roll over, so an unused allowance is simply lost, while a busy month spills into extra credits. Teams need a written rule for which pull requests get analysed and how many tests each may trigger. The FAQ’s own advice about promotion PRs is, in effect, that rule.
GitHub first
Ship’s plan table counts GitHub repositories, and its workflow examples name GitHub, Linear, Jira and Slack. The older Impact Analysis page also lists GitLab and Jenkins, but the Ship page does not mention GitLab or Bitbucket. Teams on those platforms should confirm support before starting a trial.
Agents checking agents
The pipeline Ship sketches has an AI agent testing code that another AI agent may have written, then handing the cause to a third agent to fix. That can be efficient, and it can also compound mistakes. We have written about what happens when a coding agent acts without guardrails.
Keep a human approval step on every fix, and on every test Ship adds to the permanent regression testing suite.
How to Trial Ship's Regression Testing in 15 Days
The free trial gives 3,000 one-time credits and every Starter feature. That is enough for a meaningful regression testing evaluation if you spend it deliberately rather than pointing Ship at a repository and waiting.
Budget the 3,000 trial credits
A code analysis (1,000 credits) and a crawl (500) use half the allowance before the first test runs. The remaining 1,500 credits buy, for example, 10 bug analyses (150) plus 12 pull requests at 106 credits each (1,272), with 78 credits to spare.
Plan those 12 pull requests in advance. Spending them on whatever happens to be open that week tells you less than spending them on changes you already understand.
Choose known bugs and real pull requests
Pick five to ten bugs your team has already fixed and can describe exactly, and feed them in as if they were new. Then replay pull requests from the last month where you know what broke and what did not. Known answers let you score Ship’s regression testing objectively instead of judging a polished demo.
Score every hand-off
The launch post supplies a good checklist. Can Ship reproduce a known issue with enough detail for another engineer to follow? Does its suggested testing reflect the flows a change actually affects? Are likely causes clearly separated from confirmed observations? Does the hand-off contain information your team would otherwise gather by hand?
Track the cost per useful finding
Divide the credits spent by the number of confirmed defects Ship found that your existing process missed. That one number tells you more about the value of automated regression testing than any vendor percentage, and it is the figure to compare with the cost of a contract tester or a QA hire.
Who Should Consider Ship for Regression Testing
Ship’s model of regression testing suits some teams far better than others. Size, existing test maturity and pull request volume all change the answer.
Small teams without a QA function
Unlimited users and usage-based pricing aim Ship squarely at small engineering teams that have never had a dedicated tester. For them, the $49 Starter plan with one repository and five parallel runs is cheap enough to test against real work. The realistic alternative is often no systematic regression testing at all.
Teams with a mature test suite
The FAQ says Ship “can use and maintain the testing infrastructure you already have while adding autonomous testing where coverage is missing”. Teams with established Playwright or Cypress suites should test whether Ship’s selection beats simply running their suite in parallel, and whether its generated tests duplicate ones they already have.
Where it sits in a delivery pipeline
Ship is a regression testing gate, not a pipeline. It still depends on reliable CI, reproducible preview environments and sensible branching, the unglamorous DevOps work that decides whether any automated testing pays off. Our guide to DevOps consulting costs in the UK covers what that groundwork typically costs.
If you are building that foundation, or a product that needs it, our software development team can help you design testing into the delivery process from the start.
Regression Testing With Ship: FAQs
What is Ship by ContextQA?
Ship is an AI regression testing service, launched on 5 October 2026, that tests pull requests against the deployed app, reproduces reported bugs, investigates likely root causes and turns production anomalies into new tests. ContextQA calls it an autonomous quality engineer.
Does Ship replace an existing regression testing suite?
No. ContextQA says Ship can use and maintain your existing tests while adding autonomous coverage where it is missing. In practice it chooses which tests to run for each change and writes new ones.
How much does Ship cost?
There is a 15-day free trial with 3,000 credits, a $49 a month Starter plan with 2,000 credits and a $249 a month Pro plan with 12,000. Enterprise pricing is custom. Extra credits cost $0.025 each.
Does Ship charge per user?
No. Every plan includes unlimited users, and live sessions are free under fair use, one active session per user. You pay for the work Ship does.
Is Ship an AI code reviewer?
Not in the usual sense. Code reviewers such as CodeRabbit and Greptile analyse the code in a pull request; Ship runs regression testing against the deployed application in a real browser. The two approaches complement each other.
Does Ship test mobile apps?
Not yet on real devices. Real-device mobile testing is listed as coming soon as a Pro add-on, and Enterprise customers can reserve dedicated devices.
References
Ship by ContextQA: Autonomous Quality Engineer
Introducing Ship by ContextQA: Your Autonomous Quality Engineer
TestingCatalog on X: Ship runs on pull requests
ContextQA: AI Impact Analysis for Testing
Sonar: Data Reveals Critical Verification Gap in AI Coding
Google Cloud: Announcing the 2025 DORA Report
Meta Engineering: Predictive Test Selection
Develocity Predictive Test Selection Guide