Pine Computer is a new cloud computer from Pine AI that is built for software agents rather than people. Instead of letting an AI model squint at screenshots of a desktop meant for human eyes, it gives the model a browser, files, a shell and a live screen that report what is on a page and what just changed. Developers reach it through an SDK: their product hands a Pine Computer a job, follows its progress and receives the finished work. The product launched on 9 October 2026 in private beta, following a launch post from co-founder Dylan Wang on 5 October.

The headline claim is bold. Pine says its system is two to five times faster than AI agents on conventional computers and uses about one twenty-fifth of the model cost. On UniPat AI’s public SaaS-Bench v1.1 benchmark it posts the highest checkpoint score of nine systems, 78.3%, ahead of Anthropic’s Opus 5 running in Claude Code. But it finishes fewer whole tasks than the two strongest rivals, a detail Pine itself prints next to the headline number.

This article explains what Pine Computer is, how it works and what developers get in the beta. We then read the benchmark in full, using the public results table and Pine’s own task-level data, and work out what the cost and completion figures mean for anyone thinking of building AI agents on top of it.

What Pine Computer Is

pine computer cloud environment agentic tasks b smoke alarm on a ceiling plate

Pine describes its new product as “a computer built for AI”. In practice it is three things sold as one cloud service, and the difference between them matters for anyone comparing it with other agent tools.

A cloud computer, a harness and a runtime

Pine’s own launch FAQ answers the obvious question, “Is Pine Computer a model, a harness, or a VM?”, with one word: “Yes.” In the cloud it is virtual computers plus an agent harness plus a runtime layer wired into the operating system and browser, along with local tools and remote API tools. On the developer’s side it is an SDK and an API.

That bundle is the product. A developer could assemble a browser automation service, a sandboxed virtual machine, an agent framework and a permission workflow from separate vendors. Pine Computer packages those pieces so an application can hand off a task and get back files, records or answers without exposing a general-purpose desktop to the model.

Who is behind it

Pine AI is the trading name of 19Pine Pte. Ltd. The company was founded in mid-2024 and is led by chief executive Stanley Wei. It raised a $25 million Series A in December 2025, with Fortwest Capital among the investors, according to Crowdfund Insider.

Its first products were consumer agents that handle “digital chores”: negotiating bills, cancelling subscriptions, filing complaints and chasing refunds by phone, email and web. SiliconANGLE reported in May that its plans start at $30 a month. Co-founder and chief architect Dylan Wang previously worked at Agora, the real-time audio and video company, and says that experience shaped Pine Computer: at Agora he saw how rebuilding the network underneath a call could transform its quality.

The launch at a glance

ItemDetail
ProductPine Computer, a cloud computer for AI agents
MakerPine AI (19Pine Pte. Ltd.)
Launch post5 October 2026, by Dylan Wang
Press release9 October 2026, 12:00 ET, PR Newswire
StatusPrivate beta, invite-only waitlist
AudienceDevelopers and software companies, not consumers
Entry pointSDK and API; runs only in Pine’s cloud
Headline claims2 to 5 times faster, about 1/25 of the model cost
Best benchmark result78.3% checkpoint score on SaaS-Bench v1.1

How Pine Computer Reads the Web as Structure

pine computer cloud environment agentic tasks c dumbwaiter hatch with a tray inside

The technical argument behind the launch is simple to state. Today’s computer-use agents work through machines designed for people, with a screen to look at and a mouse to point with. Pine says that wastes the model’s intelligence.

Pixels versus structure

A typical computer-use agent leans on computer vision: it takes a screenshot, works out what is on it, clicks something and takes another screenshot. Every step costs model tokens and time, and every step is a chance to misread the page. “Models have grown far more capable, but they still spend much of that capability working around a machine built for someone else,” the press release says.

Pine Computer reads a web page as structure instead: what elements are present, what can be done with them and what changed after an action. Pictures stay available for people, who can watch the live screen. The marketing site puts it as “Humans see pixels. AI reads what each thing is, and what it can do.”

Epoll for AI perception

Wang’s launch post uses a programmer’s analogy. Instead of the model repeatedly asking “what changed?”, the browser, the applications, the rendered display and the operating system send it notifications. “Think of it as epoll for AI perception,” he writes, referring to the Linux mechanism that lets a program wait for events rather than polling for them. “Changes trigger attention.”

The practical claim is about context size. Because the model does not re-read a screenshot or an accessibility tree after every step, it gets “lean, precise context”. Pine argues this is why a lower-cost model can do well on its system, and why its benchmark entry used GPT-5.6 Luna rather than a frontier model.

Where screenshots still come in

Pine’s materials are not fully consistent about other models. The press release says developers who want to bring their own model “can connect it too; it works from screenshots today”. The launch FAQ says Pine Computer currently runs Pine’s own model and GPT-5.6 Luna, and that “bringing your own model and API key is coming” once a model passes Pine’s benchmark checklist.

The two statements can both be true if outside models are possible but limited to the slower screenshot path. Either way, the structured-perception advantage Pine advertises applies today to its own preset models. Developers with an existing agent should ask which path their model would take.

The SDK: How Developers Put Pine Computer to Work

pine computer cloud environment agentic tasks d chocolate box with separate compartments

Pine Computer is not something an individual installs. It is infrastructure that other products call, and Pine is explicit that consumers will reach it only through software built on top of it.

Creating a computer for a job

The marketing site shows a short Ruby example. A product creates a computer in a chosen country, opens a browser session and gives the agent a task in plain English:

computer = pine.create_computer(
  location: { country: "US" },
  ephemeral: true
)
session = computer.create_session(browser: true)
session.agent.run("Pull last month's invoices from
  the supplier portals into the ledger.")

The product then follows the task’s progress and receives the result. “Pine runs the computer, the intelligence in it, the browser and the isolation; the developer builds the experience around it,” the release says. Pine’s site lists 26 example uses, from renewing permits on city portals to updating listings on seller portals that have no API.

Sessions, parallel work and recovery

Each Pine Computer can run several sessions, and a product can run many computers at once. Each task gets its own browser window, files, downloads and tools, so one task cannot drive another’s browser or mix up its files. During the beta each project has a concurrency allowance that Pine will raise on request.

Pine says the system is designed to recover from common failures such as a crashed browser, a dropped connection or a stopped tool. It can rebuild the affected part of the environment and keep the parts that still work, rather than restart the whole task. Sign-ins are saved per computer; sharing one browser profile across several computers is not supported yet.

Start-up, pause and resume times

The FAQ gives concrete timings. A new computer starts “in a few seconds”. One with a lot of saved state takes about 15 seconds to restore. A paused computer resumes in under a second, and infrastructure is only billed while a computer is running. Those figures matter for products that hand out many short jobs, where cold-start time can dominate the user’s wait.

The SaaS-Bench Results in Full

pine computer cloud environment agentic tasks e pinecone on a round base

The benchmark Pine leans on is SaaS-Bench v1.1 from UniPat AI, a public test of whether computer-use agents can finish real workflows in business software. It has 106 tasks across 23 deployable applications in six domains.

The nine-system table

SaaS-Bench reports two scores. The checkpoint score (CS) is the share of a task’s verification checkpoints a system passes along the way. The resolved score (RS) is the share of tasks completed end to end. The table below reproduces UniPat’s published v1.1 results.

Model and harnessSteps per taskCost per taskCSRS
GPT-5.6 Sol, open-source Browser-Use142.7~$14.969.8%17.9%
Opus 5, open-source Browser-Use162.1~$20.064.7%21.7%
Qwen 3.8 Max, open-source Browser-Use71.5~$2.531.1%8.5%
Kimi K3, open-source Browser-Use121.9~$6.856.2%17.0%
GPT-5.6 Sol, Codex250.7~$20.571.1%29.2%
Opus 5, Claude Code270.5~$26.574.3%31.1%
Qwen 3.8 Max, Qwen Code268.9~$5.165.9%23.6%
Kimi K3, Kimi-CLI284.4~$9.456.4%24.5%
Pine Computer runtime335.5~$3.678.3%27.4%

UniPat’s own summary is even-handed: under native harnesses, “Pine Computer / PCR has the highest CS (78.3%), while Opus 5 / Claude Code has the highest RS (31.1%).” The first four rows share one open-source browser harness; the last five run each model in its maker’s own agent runtime.

Checkpoints passed versus tasks finished

The chart compares the three systems Pine highlights in its release. Pine Computer leads on checkpoints by 4.0 points but trails Opus 5 with Claude Code by 3.7 points on finished tasks.

Checkpoint score versus resolved score, SaaS-Bench v1.1

GPT-5.6 Luna on Pine Computer: checkpoints 78.3%
GPT-5.6 Luna on Pine Computer: resolved 27.4%
Opus 5 with Claude Code: checkpoints 74.3%
Opus 5 with Claude Code: resolved 31.1%
GPT-5.6 Sol with Codex: checkpoints 71.1%
GPT-5.6 Sol with Codex: resolved 29.2%

In whole tasks, the gap is small but real. Pine’s own data shows 29 of 106 tasks fully resolved. At 31.1%, Opus 5 with Claude Code resolved about 33, and GPT-5.6 Sol with Codex about 31. Pine’s system makes more progress inside each task and gets stuck more often before the end.

Harness matters as much as the model

The table also supports Pine’s broader thesis that the environment around a model matters. Running the same model in its native runtime instead of the shared browser harness raised its checkpoint score in every case, by very different amounts.

Checkpoint-score gain from native harness over shared browser harness (points)

Qwen 3.8 Max: 31.1 to 65.9 +34.8
Opus 5: 64.7 to 74.3 +9.6
GPT-5.6 Sol: 69.8 to 71.1 +1.3
Kimi K3: 56.2 to 56.4 +0.2

The resolved scores moved more consistently, rising between 7.5 and 15.1 points for every model. If a harness change alone can add 35 checkpoint points to one model, the case that the computer and runtime are now as important as the model is a fair one. It also means a “best system” result says little about which component deserves the credit.

Reading the Pine Computer Benchmark Carefully

pine computer cloud environment agentic tasks f trellis with a vine climbing halfway

Pine deserves credit for publishing both scores, the run data and the caveats. Reading those materials closely raises four points that do not appear in the headline.

Cost per finished task

Model cost per task is Pine’s strongest number. At $1.02 in model tokens per task against $26.50 for Opus 5 with Claude Code, it spends about 26 times less, which is where “1/25 the model cost” comes from. Dividing each system’s cost by its resolved score gives the model cost per task actually finished.

Model cost per fully resolved task (cost per task divided by resolved score)

Opus 5 with Claude Code: $26.50 / 0.311 $85.21
GPT-5.6 Sol with Codex: $20.50 / 0.292 $70.21
Pine Computer at public table price: $3.60 / 0.274 $13.14
Pine Computer at model-token cost: $1.02 / 0.274 $3.72

Even after the lower completion rate, the cost advantage survives: about 23 times cheaper per finished task on model tokens, or 6.5 times on the public table’s price. Two cautions apply. The $1.02 excludes infrastructure, which Pine bills separately. And the public table lists Pine at about $3.60, which Pine says is the list price of its own consumer product rather than raw model tokens.

More steps, not fewer

Pine says its system is two to five times faster, but the public table has no wall-clock column, so that claim rests on what Pine calls “preliminary internal tests”. What the table does show is that the Pine Computer runtime took the most steps of any system: 335.5 per task, 24% more than Opus 5 with Claude Code and 34% more than GPT-5.6 Sol with Codex.

That is not a contradiction. A cheap model taking many small, well-informed steps can still finish sooner than a frontier model reading large screenshots. But it means the speed claim cannot be checked from public data yet, and buyers should time their own workloads.

125 attempts behind 106 results

Pine’s Hugging Face results card is unusually candid. The SaaS-Bench v1.1 batch, run on 18 and 19 September, had 125 attempts for 106 accepted outcomes. Sixteen healthcare tasks were rerun “after environment/grader fixes” and three others had grading-error retries, so 19 attempts, about 15%, were invalid or superseded. Eight timeouts remain in the results.

“This is not a first-attempt-only score,” the card says. It is fair practice to rerun tasks broken by a faulty grader, and Pine discloses it. But the comparison systems’ rerun policies are not stated in the same detail, so the like-for-like gap is uncertain by a point or two either way.

Where it struggles

The domain breakdown shows where progress and completion part company. Teamwork tasks had the highest checkpoint score, almost 87%, yet only one of 12 was fully resolved.

DomainTasksChecks passedFully resolvedResolved share
Agriculture1276.74%866.7%
Business1584.26%320.0%
Healthcare1668.98%16.3%
Media2080.85%840.0%
Software3175.85%825.8%
Teamwork1286.78%18.3%
All10678.33%2927.4%

The pattern suggests near-misses rather than failures to start: the agent gets most of the way through a collaboration or business workflow and then misses one required step. For a product that bills users for finished work, that last step is what counts. Healthcare, the domain with the reruns, is also the weakest on both measures.

Security, Logins and Human Takeover on Pine Computer

An agent that works across other companies’ websites needs credentials, and that is where most buyers’ questions will land. Pine has answered more of them than many launches do, though the answers are design claims, not audited results.

Sealed computers and your own key

“Each computer is sealed off on its own, with scoped access,” the launch FAQ says. Its saved state is “encrypted with your own key, protected by hardware in the cloud, and we can’t read it.” Pine says it will open-source this part of its implementation, “because trust has to be earned in the open”. The release adds that developers’ keys “stay with them”.

Isolation per computer limits the damage one task can do to another. It does not, on its own, stop an agent from taking a wrong action inside the accounts it has been given, which is the risk Apple flagged when it moved to restrict Full Disk Access for AI agents on macOS.

When a website asks for a password

When a task reaches a login, multi-factor prompt or approval step, Pine Computer fires a callback to the developer’s product. The user then takes over the streaming screen, types the password themselves and chooses whether the computer should remember the sign-in. Control then passes back to the agent, which continues from the same point.

This live takeover is the product’s answer to the problem every agent vendor faces: sites built to stop automation. It also keeps a human in the loop for the most sensitive moments, which is a sensible default for anything touching money or personal data.

Bot detection and CAPTCHAs

Asked whether Pine Computer will be blocked as a bot, the FAQ answers “about as often as you do on a new laptop”. It ships a browser with “a complete, PC-like fingerprint”. When a CAPTCHA appears, the AI tries it first and asks a human if it cannot get through.

That answer will reassure developers and worry some site owners, since it describes software designed to look like a person’s computer. Businesses deploying it should check the terms of the sites their agents will use.

Pine Computer Pricing and Availability

Pine has not published a price list for the developer product. It has described how charges will be built up, which is enough to estimate costs during a pilot.

What the cost includes

According to the launch FAQ, model tokens are “fully itemized” so they can be checked against official list prices. Infrastructure is billed separately, “easy to estimate from AWS and GCP list prices”, and only while computers are running. Pre-installed local software is open source and free, while remote APIs such as media production tools are billed at the provider’s rates.

That structure makes the benchmark cost a floor, not a total. A team comparing it with a self-built agent stack should add compute time, any remote API calls and the engineering time it saves.

Private beta and waitlist

The product is invite-only. Developers can request access at pinecomputer.io, and the console and docs are live for accepted projects. It runs only in Pine’s cloud; there is no self-hosted option.

Pine says it already uses the system in its own assistant and enterprise work. In one unnamed enterprise deployment it helps automate audits, and Pine says that customer’s team “now takes on 50% more work with the same people”. The release does not say how that was measured. Andrew Mackenzie, co-founder of Subliminal, which evaluated the product, said Pine had “thought about all of this and built it all in.”

How Pine Computer Compares With Other Agent Environments

Pine is entering a crowded field. Over the past few weeks we have covered AWS’s Strands Box sandbox for AI agents, InsForge’s InstaCloud serverless cloud for coding agents and Meta’s Muse agent with a dedicated VM tab. Each gives agents a separate place to work.

Three ways to give an agent a computer

Pine frames the choice as three options. The comparison below summarises Pine’s own descriptions of each, so it reflects Pine’s view of the market rather than independent testing.

QuestionAgent on your laptopBrowser automation servicePine Computer
What the agent seesScreenshots of your desktopWeb pagesPage structure and change events, with screen as backup
ScopeEverything on the machineMostly websitesBrowser, files, shell and tools
Parallel workOne pointer, one keyboardVaries by serviceMany sessions and computers
Human controlYou operate the machineUsually none built inWatch, take control, hand back
ExposureYour files, tabs and appsThe sites it visitsOnly what the task is given
Laptop must stay onYesNoNo

Where it is different

Pine Computer’s distinctive bets are the structured perception layer and the fact that it bundles the intelligence with the computer. Most sandboxes give an agent a clean machine and leave the agent logic to the developer. Pine sells the harness too, which is why its benchmark entry is a whole system rather than a model.

That bundling is a strength for a team that wants finished work from a single API call. It is a constraint for a team that has already invested in its own agent framework, especially while bring-your-own-model support is limited. Google’s work on Gemini Task mode and computer use shows the big model makers are building the same layer into their own products.

What Developers Should Check Before Building on Pine Computer

For most teams the sensible next step is a time-boxed pilot on a handful of real tasks, not a platform decision. The checklist below turns Pine’s claims into tests.

A five-point evaluation checklist

  1. Completion, not progress. Measure the share of your tasks finished end to end, since that is where Pine Computer trails on SaaS-Bench.
  2. Wall-clock time. Time the same jobs on your current agent stack, because the 2 to 5 times speed claim is not in public data.
  3. Full cost. Add infrastructure hours and remote API charges to the itemised token bill.
  4. Model path. Confirm whether your preferred model would use structured perception or screenshots.
  5. Credential handling. Walk through the login callback, the remember-me choice and what is stored, with your security team.

Questions Pine has not answered

Pine has not published prices for the developer product, a service level agreement or the concurrency limits in the beta. It has not named the enterprise audit customer or explained the 50% figure. Its OfficeVal and OSWorld results are promised but not yet out.

None of those gaps is unusual for a private beta. They are the questions to put to Pine before moving anything that touches customer accounts or money onto it. Teams weighing where agents fit in their wider plans may find our guide to AI employees and autonomous agents a useful starting point.

The Bottom Line

Pine Computer is a serious attempt to fix the part of agent software that model makers have mostly left alone: the computer the model works on. Its structured view of web pages, its live human takeover and its willingness to publish both its best and its weakest benchmark numbers set it apart from many agent launches.

The evidence so far supports a narrower claim than the marketing. On SaaS-Bench v1.1, a low-cost model on Pine Computer makes more progress per task than frontier models in their own harnesses and costs far less per task, but it finishes fewer tasks. For developers building products that hand off real work, that completion gap is the number to watch as the beta opens up.

References and Further Reading