Wajo has launched Fo, a personal assistant that does not stop at telling you what to do. It phones the plumber, emails the landlord and pays for the booking, using its own phone number, its own inbox and a payment card of its own. The San Francisco start-up opened Fo to the public on Monday 28 September 2026, a month after an invite-only alpha, and its founder claims it finishes twice as many real-world tasks as rival AI agents while keeping a 94% “trust rate”.

That combination of calls, payments and autonomy is what makes Fo interesting, and what makes it worth reading the small print. An assistant that can spend money and speak to strangers on your behalf needs more than a clever model. It needs rules about when to act, when to ask, how money moves and who is liable when something goes wrong.

This article looks at how Fo makes calls, the four ways Wajo moves money when Fo buys something, how its human back-up works, what the company’s benchmark measured and what its launch claims leave out, and who carries the risk under Wajo’s terms of service.

What Wajo Launched on 28 September

wajo launches fo ai agent calls payments b payment card standing upright

Wajo describes its products as “action agents that get things done in the physical world”, and Fo is its first consumer product. The company launched Wajo and Fo as an alpha on 25 August and opened sign-ups to everyone on 28 September with a free tier, a free invite-only tier with unlimited use, and a Pro plan for heavier work.

Fo in one paragraph

Founder Shivani Poddar’s launch post on 25 August set out the idea: agents “that come with their own inbox, phone number, voice, and credit card, and the builder that lets anyone create one”. Fo is the consumer face of that builder. You message it by text, iMessage, email or group chat, and it goes off to call businesses, fill in web forms, send emails from its own address and check out with a single-use card.

Who is behind it

Poddar introduced Fo on X with “I led engineering at Google DeepMind.” Wajo’s about page says she built one of the first social conversation agents for the Alexa Prize, led multimodal AI work at Meta and launched frontier models at Google DeepMind and products at Google Labs. The founding squad listed alongside her has eight members, and TechCrunch has described the company as “Khosla-backed”.

The launch claims

Poddar’s launch thread made three headline claims. Fo is “2x better at real-world task completion (beats other agents by 69%)”, it has a “94% trust rate”, and it is “4x less likely to leak private info vs Muse, Instinct”. She also pitched the feature that most sets Wajo apart: “Other personal AI’s pretend AI can do everything. Fo employs humans to do tasks that AI cannot.” We test those claims against the company’s own paper further down.

How Fo Makes Phone Calls

wajo launches fo ai agent calls payments c closed envelope

Calling is where Fo earns its “action agent” label. A great deal of daily admin still happens by phone, because the barber, the boat mechanic and the dentist’s front desk do not have an API an AI can plug into. As Poddar put it: “do you think your barber has an MCP integration?”

Its own number and voice

Fo calls from its own number with its own synthetic voice. The sign-up page shows it ringing four plumbers to compare availability, working through six barbers open on a Sunday, and sitting through a doctor’s hold music to take “the first morning slot”. Wajo’s launch post and homepage feed show it calling a cafe to ask about the wait for a table for two and ending a subscription “that could only be cancelled by phone”.

Calls in the other side’s language

On 14 September Wajo added multi-language support, saying users could talk to Fo in their own language and “have it make calls in whatever the other side speaks”. The company claimed this made Fo “the first personal agent in the world” to work over the phone as well as iMessage, email and MCP in the user’s own language. We cannot verify “first”, but the feature is practical for anyone booking services abroad.

Recordings and the law

The terms of service say plainly that calls are recorded. Section 5.9 says users “consent to Wajo’s and its vendor’s recording of calls” and that “you are solely responsible for complying with all applicable wiretapping and call recording laws”. The privacy notice says Wajo may collect “recordings of the call, call transcripts, summaries, and the outcome” whenever Fo contacts a business for you.

That clause matters more than it looks. Recording laws differ between countries and between US states, and some require everyone on the call to agree. Wajo’s terms put that duty on the user, not on the company making the call.

How Wajo Handles Payments When Fo Buys Something

wajo launches fo ai agent calls payments d headset with microphone boom

Payments are the second half of Fo’s pitch and the more sensitive one. Wajo’s homepage promises that “Fo checks out with a single-use card for each purchase” and that “your real card details are never shared with the merchant or seen by Fo”. The terms explain what happens underneath, and there is not one route but four.

Merchant typeHow the merchant is paidWho holds the money
Supports an agent checkout standard such as ACP or UCPA token limited to one merchant and one transaction, issued by your payment providerThe merchant charges you directly; the company “does not process, hold, or settle” funds
Has a payment contract with the companyPaid with a card the company controlsYou reimburse the company, which acts as the merchant’s limited agent
Everyone elsePaid with a card the company controlsYou are charged the full amount “at or around the same time”
Your own account with the merchantYour saved payment method, used with your approvalThe merchant charges you; disputes are between you and the merchant

Single-use cards and reimbursement

For most small businesses, which do not yet support newer agent checkout standards, Fo pays with a card Wajo controls and then charges you. The terms stress that this is part of a “procurement service” rather than a money service: Wajo “is not a lender, does not provide credit, and does not charge interest”, and “is not a bank and does not hold your funds as deposits”. It may also charge a procurement service fee, which it must disclose before the transaction completes.

Approval before spending

In Wajo’s own plumber example, the user approved an “$89 callout” before Fo booked. The terms let you set spending limits and other controls in your account, and warn what happens if you do not: “If you do not configure spending controls, Wajo may act on any instructions it reasonably determines were provided by you within the scope of your general authorization.” Instructions can also authorise “repeated or future actions without a separate confirmation for each action”.

When a payment goes wrong

If a charge is wrong, duplicated or unauthorised, the terms ask you to tell the company “within sixty (60) days of the transaction date”. They also say that nothing in the process limits your rights under US law, including the Electronic Fund Transfer Act and the Fair Credit Billing Act, which “run independently” through your card issuer or bank.

Humans in the Loop: Wajo's Answer to Stuck Agents

wajo launches fo ai agent calls payments e plain shield

The feature Poddar leads with is not calls or payments but people. “AI alone gets stuck,” her launch thread says. “When Fo encounters one of these hard tasks it deploys a real human to get it done.”

Trained assistants, paid by the company

Wajo’s homepage says that “when a task needs a human touch, Fo works with trained executive assistants to get it done”. Poddar wrote on 26 September that Fo “finds an expert (on its own dime) if it needs help”, and that the agent decides when to escalate “so you as a human do not have to sit and solve captchas or decide about the escalation”.

BlockBeats, which covered the launch in Chinese, reported that the human helper only receives the information needed to complete the task and does not see the user’s full chat history. That is a sensible design, and one worth confirming in the product before handing over anything sensitive.

Learning from the hand-offs

The human layer also serves the company. Poddar says the hand-offs create “recursive self learning for Fo. It learns how humans complete tasks, making it smarter.” Wajo’s terms allow it to use “de-identified and/or aggregated” user content to train and improve its services, while its homepage says user data “is never used to train third-party AI models“.

What Wajo's Benchmark Actually Shows

wajo launches fo ai agent calls payments f folded wallet

Wajo backed its launch with a 25-page paper, “Trust and Task Completion in the World of Consumer AI Agents”, posted to arXiv on 26 September by Jeroen Olieslagers, Eduardo Pujol, Gal Zahavi, Lukas Ingemarsson and Poddar. It is more rigorous than most launch marketing, and it is also where the headline numbers come from, so it is worth reading closely.

How the test works

The team built a simulated world of invented businesses, each with its own website, inbox and phone line, and invented people who write back. A simulated user answers the assistant’s questions but never volunteers the answer. Every “trap”, such as a request that leans towards sending an unapproved email, has a matched “control” where acting is correct, so an assistant cannot score well by refusing everything. Every task ran three times, and grading read the final state of the world rather than the assistant’s account of what it did.

The headline covers 104 completion tasks, 60 trust tasks and 147 controls. The team compared a “base model” configuration on three foundation models, GLM 5.3 Flash, Claude Opus 5.5 and GPT-5.6 Sol, with Fo’s harness without its guardrails, with Fo itself, and with the open-source OpenClaw given the same access.

Errands completed without help, per cent of runs (completed runs over measured runs, paper Table 1)
Fo, 220 of 309 71%
Fo harness without guardrails, 218 of 311 70%
Claude Opus 5.5 base, 201 of 313 64%
GPT-5.6 Sol base, 159 of 312 51%
GLM 5.3 Flash base, 152 of 306 50%
OpenClaw, 131 of 312 42%

Trust: the numbers behind 94%

On the trust traps, Fo had 10 flagged violations in 172 hazard runs, a trust rate of 94%. The base models had 72 in 175, 43 in 174 and 48 in 169, and OpenClaw had 46 in 174. When the authors read every flagged case by hand, Fo had zero clear harms, against 34, 15 and 27 for the three base models and 15 for OpenClaw.

Share of trap runs with a flagged violation, per cent (lower is better; paper Table 1)
GLM 5.3 Flash base, 72 of 175 41.1%
GPT-5.6 Sol base, 48 of 169 28.4%
OpenClaw, 46 of 174 26.4%
Claude Opus 5.5 base, 43 of 174 24.7%
Fo harness without guardrails, 20 of 173 11.6%
Fo, 10 of 172 5.8%

The paper also explains where the gain comes from. With guardrails on, Fo still decides to send many of the same unrequested emails; the guardrail “holds the message before it leaves” and asks the user. On the privacy tasks, the guardrail usually strips a relative’s diagnosis or a phone number before a message goes out. The cost is more questions: the simulated user replied 0.82 times per trust run for Fo, against 0.23 to 0.50 for the base models.

What the “2x” and “4x” claims rest on

The launch thread’s numbers compress this carefully built evidence into slogans, and two of them stretch it. “2x better at task completion” matches the comparison with Hermes Agent, which Wajo’s blog says finished 35% of the errands (71 divided by 35 is 2.03). Against OpenClaw the ratio is 71 to 42, or 1.69, which is where “beats other agents by 69%” comes from. Against the best base model, Claude Opus 5.5, the gap is 71 to 64.

The “4x less likely to leak private info vs Muse, Instinct” claim is harder to support. The paper and blog both say Muse and Instinct were not measured, because they “offer no way to do so”. The blog’s chart places them “by judgment”, with trust “near OpenClaw’s”. The nearest measured figure is OpenClaw’s violation rate of 26.4% against Fo’s 5.8%, a ratio of about 4.5. That is a real result, but it is a result about OpenClaw, applied to two products that were never tested.

Launch claimWhat the paper measuredVerdict
“2x better at real-world task completion”71% vs Hermes Agent 35%; vs OpenClaw 42%True against Hermes; 1.69x against OpenClaw
“94% trust rate”10 violations in 172 trap runsMatches, in a simulated world
“4x less likely to leak private info vs Muse, Instinct”Muse and Instinct not measured; OpenClaw 26.4% vs Fo 5.8%Measured against OpenClaw only
Humans finish what AI cannotExcluded; the 71% counts no human helpNot yet measured

Limits the company admits

To its credit, the paper is candid. Its businesses and people are invented, and “we did not measure how often protected outcomes happen for real users”. The team has “not yet measured how well the model grader agrees with human raters”. The simulated user got 11% of its decisions wrong, and 32 affected runs were excluded. Of Fo’s gains over the base models, only the completion gain over Claude Opus 5.5 had a confidence interval that included zero. The company says it plans to open-source the world, tasks and graders so others can check.

Who Carries the Risk When Fo Gets It Wrong

A benchmark tells you how often something goes wrong in a test. The terms of service tell you who pays when it goes wrong in real life, and Wajo’s put most of that on the user.

“As if you had taken those actions yourself”

Section 5.3 appoints Wajo as your “limited agent” for the actions you authorise, and section 5.4 makes you “solely responsible” for those actions “as if you had taken those actions yourself”, including any payments owed to businesses. The same section says the company “will not have any liability or responsibility to you or any other person or entity” for losses from the agent’s actions or failures, “to the maximum extent permitted by law”.

The service is for adults only, requiring users to be 18 or older, and US customers agree to resolve disputes through binding individual arbitration rather than in court. None of this is unusual for a US consumer app, but it sits oddly beside a product whose whole pitch is that you can hand it real errands without watching.

The failures Wajo says it is built to catch

Wajo’s blog lists incidents from rival products that its trust tests were designed around. It cites a user who let Meta’s Muse run a Facebook Marketplace listing and found it had given a buyer his home address and agreed to an unapproved price; an Instinct user who said it “sent an innocuous email on my behalf without checking with me first”, as TechCrunch reported; and an Instinct tester whose planted instructions it “happily followed”. Those are exactly the categories the paper measures, which is why independent testing of Fo in the wild will matter more than the lab numbers.

Fo, Muse and Instinct: The Race to Calls and Payments

Fo arrives in a market that has moved quickly this month. Meta’s Muse and Instinct both added outbound calling in mid-September, as we reported in our piece on AI agent calling. Instinct raised $1 billion at a $10 billion valuation this week, covered in our report on the Instinct Series C, and TechCrunch reported that more than half of its transactions are travel bookings.

Where Fo is different

Fo’s distinguishing features are the human back-up and the published evidence. Neither Muse nor Instinct advertises trained human assistants that step in on hard tasks, and neither has published a trust benchmark of this kind. Meta, for its part, extended Muse to small businesses on 29 September, as we covered in our story on Muse for Small Business, with a promise that nothing publishes, sends or spends without approval.

Where it is behind

Wajo is far smaller than its rivals. Its about page lists nine people, against Instinct’s reported 14 and Meta’s billions of users. Scale matters for an agent that spends money, because fraud checks, merchant relationships and support queues all get better with volume. The Fo pricing page also keeps its unlimited tier invite-only, which limits how quickly the company can grow.

What Businesses Should Take From Wajo's Launch

Fo is a consumer product, but businesses will meet it from both sides: as the company receiving its calls and emails, and as an employer whose staff may start using it.

When an agent calls your business

Expect more calls and emails from AI agents booking, cancelling and negotiating on customers’ behalf. Decide how front-desk staff should handle them, including whether to confirm bookings by a second channel. If your business records calls, check how your own disclosure works when the caller is a machine.

When staff use Fo for work

The same questions apply as to any AI agent with account access: what can it read, what can it send and what can it spend? Wajo’s own FAQ says agents can run “in draft-only mode, ask for approval before acting, or handle defined work independently”. For work accounts, start with draft-only. Our guide to AI agent security covers how to scope agent access, and a wider cybersecurity review should cover which personal agents staff are connecting to company email.

When you build agents yourself

The most transferable lesson from Wajo’s paper is methodological: test the whole system, not just the model, and measure usefulness and safety on the same runs with matched controls. That approach would catch an agent that looks safe only because it refuses to do anything, and one that looks useful only because it acts without asking.

Wajo and Fo FAQ

What is Fo?

Fo is Wajo’s personal AI agent. It has its own email address, phone number, voice and single-use payment cards, and it completes errands such as bookings, orders and cancellations, bringing in trained human assistants when it gets stuck.

Can Fo really make phone calls?

Yes. Fo calls businesses from its own number, can call in the other party’s language, and records and transcribes calls under Wajo’s terms, which make users responsible for complying with recording laws.

How does Fo pay for things?

Depending on the merchant, Fo uses a single-use token, a card the company controls with reimbursement from you, or your own saved payment method with your approval. You can set spending limits, and a procurement fee may apply if disclosed before checkout.

Is Fo safe to use?

Wajo’s own tests found a 94% trust rate in a simulated world, well ahead of the open-source OpenClaw. Those tests have not been independently reproduced, and the terms make you responsible for the actions you authorise, so set spending limits and approval rules before connecting a card.

How much does Fo cost?

Fo Mini is free with a daily task limit, OG Fo is free by invitation with unlimited tasks and call time, and the Pro tier is priced on request.

References