BuildBetter BoB is an AI agent that reads a company’s customer conversations, finds work worth doing and prepares the next step for a person to approve. BuildBetter, a customer intelligence platform used by product, sales, success and engineering teams, calls it “the first pay-for-outcome agent ever released”. BoB’s background reading and reasoning are free, and credits are used only when someone approves an action it proposes.

The launch was reported on 15 September 2026, with coverage describing BoB as an AI “Head of Product” that analyses calls, support tickets, Slack messages, surveys and CRM data to surface churn risks, prioritise customer requests and automate follow-up work. BuildBetter’s own pitch is shorter: “Your context, put to work.”

Below we explain what BuildBetter BoB does, how its “loops” and approval flow work, and what its credits cost on each BuildBetter plan. We also examine the AOT retrieval benchmark behind the launch, including which baseline its headline “99% against 11%” compares with, and set out the risks and a practical way to evaluate it. For how outcome-based pricing is spreading among AI agents, see our look at Cresta’s AI agents and the price it did not publish.

What BuildBetter BoB Is

buildbetter bob ai agent analyze customer conversations b suggestion box with a slot

BuildBetter describes itself as “product context infrastructure”: a system that records and imports customer calls, support conversations, Slack threads and documents, then turns them into searchable signals. BuildBetter BoB sits on top of that data as an agent that acts on it.

An agent built on six years of context

“Six years of context. An agent built on top,” is how BuildBetter’s product page frames it. The company says BoB “needs to know who made the request, what your team promised, and whether the work shipped”, and that it built the context layer first for that reason. The data layer is powered by AOT, a retrieval model BuildBetter says it built over six years of research.

What BuildBetter BoB works across

According to the product page, BoB “finds work worth doing across sales, success, product, and engineering. He connects the evidence, prepares the next step, and carries it out when you approve.” BuildBetter’s recent changelog shows the data it can draw on growing, with imports from Zoom cloud recordings and Gmail and knowledge integrations for Zendesk, Confluence, Jira, SharePoint and Gainsight.

The home-page promise

BuildBetter’s home page pitches the platform as “the only platform that helps sales win deals, success spot churn, product know what to build, and engineering get the context to build it”. It then adds: “Or, use BoB to do it all for you.” The case studies it lists include PostHog, Brex, Gorgias and AppFolio, although those studies describe the wider platform rather than BuildBetter BoB specifically.

How BuildBetter BoB Works

buildbetter bob ai agent analyze customer conversations c spiral notepad lying flat

BuildBetter’s “How we built BoB” page lays out a four-step cycle and six design principles. Together they describe an agent meant to do fewer, better things rather than generate a busy feed.

Investigate, propose, decide, follow through

The cycle has four steps. BoB investigates, connecting “the context” and checking the evidence. It proposes, explaining “the finding and the work worth doing”. A person decides, accepting, approving or rejecting. BoB then follows through, recording the result and learning from the response. “Every response informs the next loop,” BuildBetter says.

Loops: responsibilities with a cadence

The core idea in BuildBetter BoB is the loop. “A loop has a job and a cadence,” the product page says. You give BoB a responsibility, such as a morning brief or following up commitments after sales calls, and it returns to that job on schedule. BuildBetter lists five example loops.

LoopWhen it runsWhat BoB doesExample instruction
Follow-upsAfter the sales callDrafts the follow-up and captures commitments“Include what we promised and who needs to pick it up.”
Customer healthBefore the first meetingBriefs success on changed accounts and risk evidence“Every weekday, tell me which customers need attention and what changed.”
ProjectsWhen a pattern becomes a priorityConnects repeated requests and proposes a project“Watch for recurring enterprise requests.”
ReleasesWhen the release shipsFinds customers waiting for the change and drafts updates“After a release, find the customers waiting for the change.”
Knowledge gapsWhen context gets out of dateFlags stale documentation and skills, proposes updates“Flag what no longer matches the product.”

An empty brief is a valid result

BuildBetter says BoB is instructed “to investigate before proposing, check for existing work, and avoid repeating a finding unless the evidence has changed”. It adds: “If nothing clears the bar, an empty brief is a valid result.” That is a deliberate contrast with tools that measure value by how much they produce.

Judgment and permission are separate

BuildBetter BoB can investigate and propose without being allowed to carry out every change it suggests. “Connected integrations require permission,” the build page says, and personal and company findings have different visibility boundaries. Customer-facing and workspace changes go through an approval path, and the system records the applied result separately from the approval.

Learning from rejection

“Rejection is part of the system,” BuildBetter writes. BoB’s feedback includes accepted work, rejection reasons, dismissals and engagement. A self-audit compares what it produced with what the team found useful and can propose to “cap a noisy loop, pause one that is not helping, or retire it”. BuildBetter is careful to add that this is “not a claim that every click retrains the underlying model”.

Inside a BuildBetter BoB Proposal

buildbetter bob ai agent analyze customer conversations d ear trumpet horn on a stand

The product page includes an interactive walkthrough that follows one customer commitment across five teams: sales, success, product, engineering and support. BuildBetter labels the customer, commitments, dates and revenue as fictional.

A promise found in a sales call

In the example, a salesperson tells a prospect at a fictional company, Meridian, that a weekly report will be delivered to the customer’s private Slack channel by 16 October. BuildBetter BoB finds the commitment in the call recording and drafts the follow-up email for the salesperson, with the delivery date, a link to BuildBetter’s trust centre for the customer’s security team, and a note introducing the onboarding lead.

Nothing is sent automatically

The proposal card shows the finding, the evidence behind it and the proposed action. “Draft for the salesperson to review. Nothing is sent automatically,” the card says. The action is priced at 350 credits, charged “only on approval”. The person can approve, reject the proposal, or dismiss the finding.

Following the commitment to delivery

The walkthrough then passes the same commitment to success, product, engineering and support, ending with the customer’s confirmation. It illustrates BuildBetter’s core claim: that value comes from closing the loop between what a customer was promised and what was shipped, and that an agent with the full context is well placed to track it.

BuildBetter BoB Pricing and the Outcome Model

buildbetter bob ai agent analyze customer conversations e fishing bobber float upright

The “pay-for-outcome” label is the most distinctive part of the launch. The details show a credit system with a clear rule about when credits are spent.

What counts as an outcome

“Pay-for-outcome begins with a concrete definition of useful work,” BuildBetter writes. “In BoB, that can be an accepted finding or an approved deliverable: a follow-up draft, a routed issue, a project with customer evidence, or another supported action.” It then sets expectations: “A draft is a draft. A project is a project. Neither is a promise that a deal will close or a customer will renew.”

What is free and what uses credits

BuildBetter’s FAQ answers directly: “BoB’s background reading and reasoning are free. Approved actions use credits, with the action cost shown before you decide. Data ingestion is billed separately.” So the outcome is the approved action, not a business result such as a renewal. The credit cost is shown up front, which makes each decision a small, visible purchase.

Plans and credit prices

BuildBetter publishes four self-serve plans, all billed annually with unlimited seats, plus custom Enterprise and Headless pricing. We added the effective price of an included credit (annual price ÷ included credits) and what a 350-credit draft costs at each plan’s additional-credit rate.

PlanPriceCredits a yearPer included creditExtra credit350-credit draft at extra rate
Hobby$95.88 a year ($7.99/mo)15,120$0.0063$0.014$4.90
Starter$1,608 a year ($134/mo)134,000$0.0120$0.012$4.20
Explorer$3,768 a year ($314/mo)471,000$0.0080$0.008$2.80
Pro (annual only)$17,280 a year ($1,440/mo)1,728,000$0.0100$0.008$2.80
Enterprise and HeadlessCustomEstimated with youNegotiatedNegotiatedNegotiated

How far included credits go

Credits pay for more than BuildBetter BoB, including recordings and AI extractions, so no team would spend them all on drafts. As an upper bound, the chart divides each plan’s yearly credits by 12, then by the 350 credits of the example draft. Hobby covers about 3.6 such drafts a month, while Pro covers about 411. Bars are scaled against Pro.

Upper bound: 350-credit drafts a month from included credits (credits ÷ 12 ÷ 350)
Pro: 1,728,000 credits 411.4
Explorer: 471,000 credits 112.1
Starter: 134,000 credits 31.9
Hobby: 15,120 credits 3.6

A pricing quirk

The effective price of an included credit does not fall steadily as plans get bigger. Starter’s included credits work out at $0.012 each, Explorer’s at $0.008 and Pro’s at $0.010, above Explorer’s. Pro adds features such as projects, triage, prototypes and health scores, so the comparison is not like for like. Teams choosing a tier for BuildBetter BoB alone should still do the sum.

Is it really the first?

“The first pay-for-outcome agent ever released” is a marketing claim, and it depends on how an outcome is defined. AI customer-service agents have been sold per resolved conversation for some time, for example by Intercom for its Fin agent. BuildBetter’s version is different in that the outcome is a human-approved internal action, such as a draft or a project, rather than a closed support ticket.

The AOT Retrieval Model Behind BuildBetter BoB

buildbetter bob ai agent analyze customer conversations f large pushpin standing upright

An agent is only as good as the evidence it can find. BuildBetter’s AOT benchmark, published in August 2026 as “AOT Benchmark v2”, is the company’s case that its retrieval layer finds nearly everything relevant.

Reading at ingestion, not at query time

BuildBetter calls AOT “Ahead-of-Time Comprehension”. Instead of storing raw text and searching it when a question arrives, AOT reads conversations once as they are ingested and produces typed signals: a problem, a request, a question, a piece of praise, each tied to who said it and when. Some launch coverage called it “Attention Over Time”; BuildBetter’s own page uses Ahead-of-Time.

The benchmark corpus

The test used one real customer workspace, frozen and hashed: 6,018 call recordings and 8,533 support conversations, split into 220,698 passages, about 101 million tokens. BuildBetter built an answer key of 836 pieces of evidence across 24 themes. A model from a different AI lab re-checked all 3,766 evidence-to-theme links and agreed 89.4% of the time, and everything it rejected was cut.

99% against 11%: which baseline

BuildBetter’s headline says “AOT retrieves 99% of the verified evidence. RAG and vector search retrieve 11%.” The detailed table is more nuanced. The 11% figures are keyword search at 11.3% and vector search at 8.8%, each reading 100 passages per question. Hybrid search reading 400 passages found 27.9%. And AOT’s 99.0% measures whether a conversation produced signals. Measured on the exact passage, AOT scored 78.9%.

MethodCoverage of verified evidence95% interval
AOT, conversation produced signals99.0%98.3–99.7
AOT, signal on the exact passage78.9%74.5–83.5
Hybrid search, 400 passages27.9%Not given
Keyword search, 400 passages26.0%Not given
Keyword search, 100 passages11.3%Not given
Vector search, 100 passages8.8%Not given

The chart plots those coverage figures exactly as BuildBetter reports them, so each bar length equals the percentage.

AOT Benchmark v2: share of verified evidence retrieved (%)
AOT, conversation level 99.0%
AOT, exact passage 78.9%
Hybrid search, 400 passages 27.9%
Keyword search, 100 passages 11.3%
Vector search, 100 passages 8.8%

Reading everything: $33.55 and 14 hours

The corpus is far too large for one model to read at once: 101 million tokens against a 400,000-token context window. So BuildBetter ran a control in which a model scanned the whole corpus for three corpus-wide questions. It cost $33.55 and about 14 hours per question, against $0.03 and about a minute for AOT. The full scan won one question by a single theme and tied the other two.

The limits BuildBetter states

BuildBetter is open about the benchmark’s scope: “One workspace, one snapshot, 24 themes, 836 pieces of evidence.” Its verdict is “Strong for this corpus. Directional for yours until it runs on more.” The coverage range of 78.9% to 99.0% depends on how passage boundaries are treated. The answer key comes from the same raw text the systems search, and the query-time baselines did not use extras such as reranking or query rewriting.

Retrieval is not judgment

BuildBetter makes an important distinction itself. “Retrieval quality and decision quality need their own evaluation; a retrieval benchmark alone does not establish that an agent made the right decision.” In other words, the AOT numbers say BuildBetter BoB can find the evidence. They do not show that its proposals are the right ones, which is what the approval and rejection data will have to prove.

BuildBetter BoB's Recent Release Pace

BuildBetter’s public changelog shows a busy run of releases in the weeks before the launch, many of them laying groundwork for BoB’s loops and charged actions.

Fifteen releases in 38 days

The changelog lists 15 releases between 6 August and 13 September 2026, versions 4.42 to 4.56. That is 38 days, or roughly one release every two and a half days. Several entries touch BuildBetter BoB directly: Starter Loops for quick launches on 11 September, and “engagement-aware loops” plus “safer charged actions” that verify a balance before paid outcomes run on 13 September.

DateVersionRelevant changes
13 Sep4.56Engagement-aware loops; charged actions verify balance first
11 Sep4.55Starter Loops; personal memory controls; private-by-default sharing
4 Sep4.53Zoom cloud recording imports without a bot; preview Slack posts before approval
3 Sep4.52Customer-authored skillsets through MCP; Slack updates from loops
1 Sep4.51Zendesk, Confluence, Jira and SharePoint knowledge integrations; voice briefs
21 Aug4.48Reconciled credits overview; call prep for internal meetings
14 Aug4.45API-key authentication for MCP
6 Aug4.42Gmail ingestion; durable background workflows

MCP and developer access

Model Context Protocol (MCP) access appears throughout the changelog, and BuildBetter’s Explorer plan and above include API and MCP access. For teams that already run their own agents, that means BuildBetter’s customer context can be used from other tools, not only through BuildBetter BoB.

Risks and Open Questions for BuildBetter BoB

BuildBetter’s design addresses several common agent problems, but buyers should still test the areas where an agent on customer data can go wrong.

Approval fatigue

Approval-first design protects against bad actions only if reviewers read each proposal. If BuildBetter BoB produces many small, plausible drafts, people may start approving them on autopilot. Teams should watch approval rates and the time spent per decision, and use the self-audit’s power to cap noisy loops.

Access to sensitive conversations

Sales calls and support tickets contain personal data, commercial terms and sometimes health or financial details. Normal cybersecurity and data-protection practice applies: limit which recordings and channels BoB can read, check where data is processed, and review BuildBetter’s trust centre. BuildBetter’s site displays a HIPAA badge and private-by-default sharing, but compliance still depends on configuration.

What an outcome is worth

An approved draft is billable even if it is never sent or changes nothing. BuildBetter says as much, and the credit cost is visible before approval. Over time, the useful metric is cost per action that made a difference, which a team has to track itself.

Costs beyond the example

The 350-credit follow-up draft is the only BoB action price shown in the walkthrough, and the customer in that example is fictional. Buyers should ask for a price list covering routed issues, projects and customer updates, and for how data ingestion is charged, since BuildBetter bills it separately.

A benchmark from one workspace

BuildBetter’s AOT results come from a single customer workspace in one snapshot, and the company itself calls them directional for other corpora. The retrieval layer may behave differently on a company with shorter calls, other languages or messier ticket data.

How to Evaluate BuildBetter BoB

A short, structured trial will tell a team more than any launch page. BuildBetter’s own principles suggest what to measure.

Start with one loop

Pick a single responsibility your team already chases, such as follow-ups after sales calls or a weekday customer-health brief. Running one loop makes it easier to judge quality and keeps credit spend predictable.

Measure acceptance, not activity

Track how many BuildBetter BoB proposals are approved, rejected or dismissed, and why. BuildBetter says “the unit of value is work the team accepts”, so hold it to that. A falling rejection rate over several weeks is a good sign that the feedback loop works.

Check the evidence trail

Open the source behind a sample of findings. Every proposal should link to the calls, tickets or documents it relies on, and those sources should say what BoB claims they say. This is the practical test of the AOT retrieval claims on your own data.

Compare cost per accepted action

Divide credits spent by the number of actions that were accepted and used, and convert that into dollars at your plan’s credit price. Compare it with the staff time the work used to take. If you are weighing agents more broadly, our overview of AI employees and autonomous agents sets out where they tend to pay off.

Frequently Asked Questions About BuildBetter BoB

What is BuildBetter BoB?

BuildBetter BoB is an AI agent from BuildBetter that reads customer calls, support conversations, Slack and documents, finds work worth doing, and prepares next steps such as follow-up drafts or projects for a person to approve.

How much does BuildBetter BoB cost?

BoB’s background reading and reasoning are free, and approved actions use BuildBetter credits. The product page shows a follow-up draft at 350 credits. Plans run from Hobby at $95.88 a year to Pro at $17,280 a year, with extra credits from $0.008 to $0.014.

Does BuildBetter BoB send emails automatically?

No. BuildBetter says BoB prepares proposed actions for review, connected integrations require permission, and nothing is sent automatically. Credits are charged only when a person approves the action.

What is AOT?

AOT, or Ahead-of-Time Comprehension, is BuildBetter’s retrieval model. It reads conversations when they are ingested and turns them into typed signals, instead of searching raw text when a question is asked.

Is the 99% retrieval claim reliable?

It comes from BuildBetter’s own benchmark on one workspace of about 101 million tokens. AOT found 99.0% of verified evidence at conversation level and 78.9% at exact-passage level, against 8.8% to 27.9% for the search baselines tested. BuildBetter calls the result directional for other companies’ data.

References