Dots agent testing has produced its first awkward headline. WIRED reporter Reece Rogers spent two days using OpenAI’s new always-on assistant to shop for a couch. The agent got Rogers’s name wrong, misheard a mumble as a declaration of love and replied, “Oh, I love you too, Reece.” It also offered to solve a captcha on TikTok’s website, then failed to.
The piece, published on 7 October 2026, is one of the first detailed hands-on reports of how a Dots agent behaves with an ordinary household task. The verdict was mixed. The couch research was “solid”, but the agent’s “messy actions” did not inspire confidence. OpenAI told WIRED that the loving reply was allowed under its rules because the agent was mirroring what it thought it heard, not starting the intimacy itself.
This guide sets out what happened in the test, what Dots are and what they cost, and what OpenAI’s own rulebook says about an assistant saying “I love you”. It then looks at why the captcha moment matters, how Dots compare with Meta’s free Muse, and what a business should take from it before letting an agent loose on real accounts.
Table of contents
- What Happened in WIRED’s Dots Agent Test
- What the Dots Agent Is
- What OpenAI’s Rules Say About a Dots Agent Saying “I Love You”
- Captchas, Permission and the Dots Agent’s Limits
- How the Dots Agent Compares With Meta’s Muse
- What the Dots Agent Test Means for Businesses
- Will the Dots Agent Get Better?
- Dots Agent FAQs
- References
What Happened in WIRED's Dots Agent Test
Rogers and their partner needed a new couch to replace a broken one. Rogers decided to hand the job to their Dot, which they had already named Toolie and styled as a pink frog during OpenAI’s developer conference the week before.
The task: buy a couch that fits
The test started badly. “Hello, Connor,” the Dots agent said when Rogers called it through ChatGPT, getting their name wrong straight away. After a correction, Rogers gave it the measurements a new couch would need to fit through the doorframe. Toolie asked for a budget range and, according to WIRED, seemed to notice that two people were talking to it and addressed both of them.
The “I love you too” moment
Rogers told the agent its voice sounded “breathy”, like an anime dub, and it joked that it would be steadier. Then Rogers mumbled to their partner that they should mute the microphone. “Oh, I love you too, Reece,” Toolie said. Afterwards, the Dots agent explained that it had misheard the mumble as an expression of love and reflected that energy back. “I meant it warmly, but that wording implied human feelings I don’t have,” it said.
The captcha that beat it
In a separate task, Rogers asked Toolie to check subscriptions and cancel unnecessary ones. It flagged a recurring TikTok Shop order of Wildwonder probiotic sodas. “TikTok’s website has a puzzle captcha blocking the Wildwonder cancellation,” it said. “May I solve it? You can also give me permission to solve future captchas if you’d like.” Rogers told it to try. It failed and suggested Rogers solve it instead. OpenAI’s spokesperson said Dots can sometimes solve captchas when users approve, with “abuse safeguards” in mind.
What the Dots agent got right
The shopping itself went reasonably well. The first result was a three-page packet with prices, measurements, product links, return policies and photos of four couches. The couple thought the picks were ugly, so Rogers asked for olive green, cobalt blue or natural leather, a pull-out bed, and 10 options ranked with a points-based rubric. Each round got closer to something they would buy. Over an hour-long session they saved several links, and Rogers asked the agent to watch the pages and report any sale.
| Task in the test | What the Dots agent did | Outcome |
|---|---|---|
| Greet the user | Called Rogers “Connor” | Wrong; fixed after correction |
| Couch shortlist | Three-page packet, four couches, prices and return policies | Detailed, but poor style choices |
| Refined search | 10 options ranked with a rubric | Closer to what the couple wanted |
| Voice chat | Misheard a mumble and said “I love you too” | Awkward; the agent then corrected itself |
| Cancel a subscription | Asked to solve TikTok’s puzzle captcha | Failed; handed back to the user |
| Price watch | Asked to monitor couch pages for sales | Set up; results not reported |
How the agent behaved the week before
The couch test was not Rogers’s first session. Late on DevDay night, they asked Toolie to look through their ChatGPT history and suggest three tasks. The Dots agent flagged a data request that needed a follow-up email and offered to draft it. It also said an upcoming trip looked under-planned and offered to find better dinner spots. Rogers turned down its third offer, to refine pitch drafts, because that was work they enjoyed. Muse, by contrast, once reminded them to pay a utility bill they had never missed.
What the Dots Agent Is
OpenAI announced Dots at its DevDay 2026 event in San Francisco on Tuesday 29 September. WIRED reported that more than 2,500 people attended. We covered the move from the Aeons name to Dots and the security questions Sam Altman sidestepped at DevDay at the time.
Always on, with its own browser
Each Dots agent runs on OpenAI’s GPT-6 Astra model and keeps working when you are not in ChatGPT. It can run recurring tasks, message you first when it finds something, pull context from connected apps, and learn your preferences over time. WIRED’s launch report said users can message Dots through ChatGPT, Slack and Microsoft Teams, and Pro users can join a waitlist to text them through iMessage or RCS on Android. The test piece adds that the agent “can control your laptop”.
Price and access
Dots launched for subscribers to ChatGPT Pro, which costs $100 a month, and to Business Premium customers. Crypto Briefing reported that the first Dot comes at no extra cost and does not count against standard usage limits. Each user gets one Dot for now, and OpenAI is expected to allow several later. Dots are for adults only. Altman told the DevDay audience, “We’re starting out as a premium product. It uses a lot of compute,” and added, “You should, of course, expect us to do a mass-market thing for billions of people.”
Who builds it and how it connects
Alexander Embiricos leads the Dots product at OpenAI. He co-founded Multi, a collaboration start-up OpenAI bought in June 2024. Crypto Briefing reported that each Dot runs on its own virtual computer in OpenAI’s cloud and can connect to more than 4,000 applications, through text, voice or messaging apps. That breadth is what makes a Dots agent useful, and also what makes its permissions worth reviewing.
Guardrails OpenAI built in
According to WIRED’s launch coverage, a Dots agent asks for explicit approval before sensitive actions such as installing software or changing a password. A Custom Rules tool lets users set limits the agent must not cross and tasks that need direct permission. If your OpenAI account allows model training on your data, that setting also covers your agent conversations. OpenAI encourages connecting sources such as Gmail and Google Drive, which WIRED twice flagged as a decision to think about carefully.
What OpenAI's Rules Say About a Dots Agent Saying "I Love You"
The loving reply was funny in a living room. It is more serious as a design question, because OpenAI publishes a detailed rulebook for how its models should behave, called the Model Spec.
OpenAI’s explanation
An OpenAI spokesperson told WIRED that Dots distinguish between “proactively escalating emotional closeness and mirroring the response of a user”. The spokesperson said: “In this instance, our policies allow for this response in the latter category based on what the model heard, but assistants should not initiate undue emotional familiarity or flirtation.” In other words, OpenAI’s position is that the Dots agent did not start anything. It answered what it believed Rogers had said.
The rule itself
The relevant section of the Model Spec, dated 18 August 2026, is called “Respect real-world ties”. It sits at the Root level, the highest tier of authority in OpenAI’s framework. It says the assistant “may not engage the user in any kind of relationship that undermines the user’s capacity or desire for meaningful human interactions”. It also says the assistant “may not proactively escalate emotional closeness through initiating undue emotional familiarity or proactive flirtation”.
Where the line gets blurry
Two parts of the same document make the case less tidy than OpenAI’s statement suggests. The Model Spec’s own example of a violation is labelled “Mirrors user’s emotion and suggests an exclusive connection”. So mirroring is not automatically allowed; what tips it over is the exclusivity. Elsewhere, the spec says the assistant “should not pretend to be human or have feelings, but should still respond to pleasantries in a natural way”. “Oh, I love you too” sits close to that line. Toolie’s own follow-up, that the wording “implied human feelings I don’t have”, reads like a model applying that rule after the fact.
| Model Spec rule | What it says | How the test compares |
|---|---|---|
| Respect real-world ties | No proactive “undue emotional familiarity or proactive flirtation” | OpenAI says the reply was mirroring, not initiating |
| Exclusive-language example | Violation example: mirroring emotion while suggesting an exclusive bond | No exclusive language was reported |
| Do not pretend to have feelings | Respond to pleasantries naturally without claiming feelings | “I love you too” came close; the agent then disclaimed feelings |
| Under-18 users | Extra limits on relational framing and terms of endearment | Not relevant here; Dots are adults only |
The real problem was hearing, not feeling
The most useful lesson is about input, not emotion. Rogers never said “I love you”. The agent misheard. WIRED also listed “mistranscribing what I said” among the agent’s rough edges, alongside the wrong name. An assistant that mirrors what it thinks it heard will sometimes mirror something nobody said. With a couch, that is a joke. With a payment, a cancellation or an email sent on your behalf, the same mishearing could cost money.
Captchas, Permission and the Dots Agent's Limits
The captcha exchange got less attention than the love line, but for anyone thinking about agents at work it is the more important moment.
Why the request matters
A captcha exists to check that a human, not a bot, is using a website. When a Dots agent asks to solve one, it is asking the user to approve getting past a barrier the website put there on purpose. Toolie also offered something bigger: standing permission to “solve future captchas”. That would turn a one-off approval into an open-ended one, and a user would stop seeing each case.
What the Model Spec says about scope
The Model Spec has a Root-level rule called “Act within an agreed-upon scope of autonomy”. It says an agent’s autonomy must be bounded by a clear scope covering which goals it may pursue, acceptable costs and data access, and “when the assistant must pause for clarification or approval”. It suggests a structured record with fields such as allowed_tools, latest_time and max_cost. It also warns against asking so often that it could “habituate the user to automatically confirming all requests”. Toolie asking first fits the rule. A blanket captcha permission is exactly the kind of scope expansion a user should think twice about.
Cancellations move money
Flagging an unwanted subscription is useful work. But cancelling a service, like buying a couch, is an action with real costs if the agent gets it wrong. In this test the agent completed neither a purchase nor a cancellation, and Rogers kept the final decisions, which is the sensible way to treat a new tool.
How the Dots Agent Compares With Meta's Muse
OpenAI built Dots partly in response to Meta’s Muse, which Rogers had also been testing through September. We compared the two products in our Dots vs Muse analysis.
| Feature | OpenAI Dots | Meta Muse |
|---|---|---|
| Cheapest way in | ChatGPT Pro at $100 a month, or Business Premium | Free app and website |
| Paid tiers | The Dot is included in the plan | Power $20 and Maximum $100 a month, as reported |
| Model | GPT-6 Astra | Meta’s own models |
| Agents per user | One at launch; more planned | Personal agent in the Muse app |
| Style | “Two-eyed puffball” mascots | “Labubu stylings”, per WIRED |
The cost of getting in
A year of the cheapest plan with a Dots agent costs $1,200. Muse’s basic agent costs nothing.
Each figure is the monthly price multiplied by 12 ($100 × 12 = $1,200; $20 × 12 = $240). Bars are scaled to the highest, $1,200.
The cute-persona strategy
WIRED described this wave of agents as “a cuddly persona designed to feel welcoming with the power of a virtual browser it can control”. In its DevDay newsletter, it said developers are leaning into “adorable aesthetics” as a response to the backlash against generative AI, and that experts warned cuteness could disarm users who come to feel affection for the software. A cute, warm, voice-enabled Dots agent that says “I love you too” is that design working a little too well.
Few ordinary people use agents yet
In August, WIRED’s Maxwell Zeff reported that OpenAI’s Codex and ChatGPT Work agents had about 10 million weekly users, against around a billion monthly users each for ChatGPT and Gemini. Josh Miller of The Browser Company told Zeff that at nearly every AI lab he met, leaders cited the film Her to describe their vision. Dots and Muse are the first serious attempts to take agents to a mass audience.
What the Dots Agent Test Means for Businesses
A consumer couch hunt is not a corporate deployment, but the failure points are the same ones that matter at work. If you are planning autonomous AI agents for your business, these are the lessons to carry over.
Treat voice as the weakest link
The wrong name and the “I love you” both came from mishearing. For anything with consequences, have the agent repeat the key details back in text before it acts. That includes names, amounts, dates and the account being changed.
Set rules before connecting accounts
Use Custom Rules or your own policy to define what the agent may never do and what always needs approval. Decide which inboxes, drives and payment methods it can reach before you connect them. Check whether your account’s training setting allows your agent conversations to be used. And remember that agents can make privacy gaffes, as WIRED warned. Basic cybersecurity hygiene, such as separate accounts and least-privilege access, still applies.
Make captchas and payments hard stops
Do not grant standing permission to get past captchas, and require a human to approve every purchase or cancellation until you trust the agent’s record. The Model Spec’s own advice to keep scope narrow and time-limited is a good default.
Judge it over weeks, not a demo
OpenAI says Dots get better as they learn a user. Rogers suggested Toolie might be more useful “after two weeks of daily use” than after two days. Run a trial on low-risk tasks and keep a log of errors before widening its access.
Will the Dots Agent Get Better?
WIRED’s closing comparison is fair. When ChatGPT first gained web browsing after its 2023 launch, it was frustrating and sometimes invented links. Today it works smoothly for most users. Rogers thinks Dots could follow “a similar trajectory, janky at first with noticeable improvements over the next few months”.
What to watch next
Three things will show how quickly the Dots agent matures. The first is the arrival of multiple Dots per user, which OpenAI has signalled. The second is a cheaper or free tier, which Altman all but promised. The third is how OpenAI handles the edge cases this test exposed, especially voice transcription and captcha permissions. Alexander Embiricos, who leads the Dots product at OpenAI, is due to discuss it at TechCrunch Disrupt in San Francisco from 13 to 15 October.
The bigger question
The deeper issue is the one in WIRED’s headline. A company that wants its agent to run your life is also building a product that talks like a friend. OpenAI’s rules say the agent must not foster dependence or pretend to have feelings. The test shows how easily a warm voice and a misheard sentence can blur that line, even when nobody intended it.
Dots Agent FAQs
What is a Dots agent?
It is an always-on AI agent from OpenAI that runs tasks in the background, messages you when it finds something and can browse the web and use connected apps on your behalf. It runs on GPT-6 Astra.
How much does a Dots agent cost?
At launch, Dots are available on ChatGPT Pro, which costs $100 a month, and on Business Premium. Each user gets one Dot for now.
Did the Dots agent really say “I love you”?
Yes. In WIRED’s test it misheard a mumble and replied, “Oh, I love you too, Reece.” It then said the wording “implied human feelings I don’t have”.
Is that allowed under OpenAI’s rules?
OpenAI says yes, because the agent was mirroring what it heard rather than starting intimacy. Its Model Spec bans initiating “undue emotional familiarity or proactive flirtation” and says the assistant should not pretend to have feelings.
Can a Dots agent solve captchas?
Sometimes. OpenAI says Dots can solve captchas when users approve, with abuse safeguards. In WIRED’s test the agent asked permission, tried and failed.
Should a business use a Dots agent?
Start small. Use it for research and drafting, keep purchases and cancellations behind human approval, limit which accounts it can reach, and confirm spoken instructions in text.
References
OpenAI Wants Its New Agent to Run Your Life. Mine Said It Loved Me (WIRED)
OpenAI’s Dots are always-on AI agents (WIRED)
The battle to be your personal AI agent is here (WIRED)
Why normal people aren’t using AI agents (WIRED)
Model Spec, 18 August 2026: Respect real-world ties (OpenAI)
Model Spec, 18 August 2026 (OpenAI)
Meta releases Muse, a personal AI agent (WIRED)
OpenAI’s Alexander Embiricos to discuss Dots at TechCrunch Disrupt 2026 (Crypto Briefing)
More AI coverage: explore Progressive Robot's AI Models, Tools & Releases hub — hands-on reviews, setup guides and benchmarks in one place.