Hybrid mode is the name Perplexity has given to an unreleased feature of Perplexity Computer on Mac that lets the cloud agent hand selected subtasks to a model running on the user’s own machine. TestingCatalog surfaced it on 31 August 2026, describing a setting that “lets Computer continue orchestrating in the cloud while automatically delegating suitable subtasks to a model running locally”, and a companion local model called Privacy Gate that checks data for personal information before anything leaves the Mac.

The stated purpose is cost. Perplexity Computer meters everything it does in credits, and the credit burn has been the loudest complaint since the product launched in February. Delegating cheap subtasks to a model the user already paid for, on hardware the user already owns, is the most direct answer to that complaint Perplexity has produced for Mac users, who were left out of last week’s Nvidia-only Portable Computer launch.

This article explains what Hybrid mode does and does not move off the cloud, the three downloadable models behind it and the RAM each needs, which Macs qualify, how Privacy Gate differs from the routing design Perplexity showed at Computex in June, the credit arithmetic the feature is aimed at, and what a UK business should make of a feature that is still hidden behind a flag. Every figure is traceable to a link in the References section.

What Hybrid Mode Does Inside Perplexity Computer

perplexity hybrid mode mac local models subtasks b solid cloud shape

Perplexity Computer is the agent product that Perplexity launched on 25 February 2026, first for Max subscribers. It takes a goal, breaks it into subtasks, and routes each one to whichever of around twenty models it judges best, all of it running in the cloud. We covered the product’s workflow uses when it appeared in Perplexity Computer: 7 Powerful AI Workflow Uses. Hybrid mode changes where some of those subtasks run, not how the product plans them.

The cloud still orchestrates

The most important detail in the TestingCatalog report is that orchestration stays in the cloud. Hybrid mode does not turn the Mac into the agent; it turns the Mac into one more worker the cloud agent can call. The planner, the tool selection, the sandbox and the connectors remain on Perplexity’s side, and a local model picks up the subtasks that the orchestrator decides are suitable for it. That is a narrower and more conservative design than Portable Computer, where the whole harness runs on the user’s GPU.

Subtasks are the unit of delegation

Perplexity’s release notes of 24 August 2026 record that GPT-5.6 Terra became the default model for Computer’s subagents. Those subagents are the workers that handle the individual steps inside a longer job, and each one consumes credits. Hybrid mode inserts a local model into that list of available workers, so some steps that would have gone to a cloud subagent go to the Mac instead. Which steps qualify is not spelled out in the report, and “suitable subtasks” is the only phrase Perplexity’s interface offers.

Where it sits in the Computer lineage

Hybrid mode is the fourth distinct shape Perplexity Computer has taken in six months. The table below places it against the other three.

ProductDateWhere the agent runsHardwareStatus
Computer25 Feb 2026Entirely in the cloudAny browserLive, Pro and Max
Personal ComputerMar to May 2026Cloud agent with local file, app and browser accessAny Mac on macOS 14 or laterLive, Pro and Max
Portable Computer25 Aug 2026Entirely on the user’s GPU, cloud escalation optionalNvidia RTX with 24GB VRAM or DGX SparkLive on Linux, Windows due September
Hybrid mode31 Aug 2026 (leak)Cloud orchestrator, local model for suitable subtasksMac with 16GB or 32GB RAMHidden, no release date

The Three Local Models Behind Hybrid Mode

perplexity hybrid mode mac local models subtasks c solid mailbox raised flag v2

TestingCatalog found three downloadable models attached to the Hybrid mode setting, each with a download size and a memory requirement. None of them is a frontier model, and none is meant to be; they are the workers for the subtasks the cloud decides not to keep.

Perplexity’s own model

The first option is a proprietary Perplexity model of roughly 19GB that requires a Mac with 32GB of RAM. Perplexity has form here: Portable Computer shipped with PPLX 27B, a post-trained version of Alibaba’s Qwen 3.8 27B, so a Perplexity-branded local model of similar scale is consistent with what the company has already released for Nvidia hardware. Whether the Mac model is the same weights repackaged for Apple silicon is not stated.

Qwen 32B

The second option is a Qwen 32B model at roughly 17.4GB, also requiring 32GB of RAM. Alibaba’s Qwen family is now the default choice for anyone shipping a local agent, and a 32B model at 17.4GB implies an aggressively quantised build, since a 32-billion-parameter large language model at full precision would need several times that space. Quantisation is what makes a model of this size fit on a consumer Mac at all.

A Gemma-based model for 16GB Macs

The third option is a smaller Gemma-based model of roughly 5.6GB that runs on Macs with 16GB of RAM. This is the entry that matters for most of the installed base, because 16GB is the starting configuration of Apple’s current Mac mini and of most laptops in the fleet of a typical UK business. It is also, by size, a much weaker worker than the two 32GB options, and Hybrid mode on a 16GB Mac will be able to take on correspondingly fewer subtasks.

Local model optionDownload sizeRAM requiredLikely role
Proprietary Perplexity modelAbout 19GB32GBStrongest local worker, Perplexity-tuned for agent steps
Qwen 32BAbout 17.4GB32GBOpen-weight alternative of similar scale
Gemma-based modelAbout 5.6GB16GBEntry option for base-spec Macs, lighter subtasks only

The chart below scales the three downloads against the largest of them, which makes clear how much smaller the 16GB option is: the Gemma-based model is under a third the size of the Perplexity model.

Hybrid mode local model download sizes, scaled to the largest (19GB = 100%)
Proprietary Perplexity model 19GB
Qwen 32B 17.4GB
Gemma-based model 5.6GB

Which Macs Can Run Hybrid Mode

perplexity hybrid mode mac local models subtasks d solid safe box round dial

The memory requirements in the leak map directly onto Apple’s current line-up, and the mapping is less generous than it first appears. Hybrid mode is a feature of the Mac app, which requires macOS 14 Sonoma or later, but the local models add a memory floor that the app itself never had.

The 16GB floor and the 32GB tier

Apple’s Mac mini specifications list three M6 configurations, two of which start at 16GB of unified memory and can be configured to 24GB or 32GB, and a third that starts at 24GB with a 32GB option. The M5 Pro Mac mini starts at 24GB and goes to 48GB or 64GB. Only the 16GB Gemma-based option works on the base machines; the two stronger models need a 32GB configuration or better, and a 24GB Mac, which Apple sells in volume, qualifies for neither of them under the requirements TestingCatalog reported.

Why a Mac mini is the reference machine

Perplexity has said since the Personal Computer launch that “running Personal Computer on a Mac mini creates the best experience because it allows the agent to run continuously”. Hybrid mode strengthens that argument, because a local model that is always loaded on an always-on desktop is exactly the worker a long-running cloud agent wants to call. A laptop that sleeps in a bag is a worker that disappears mid-job.

Mac mini configuration (Apple specs)Unified memoryGemma-based (16GB)Qwen 32B or Perplexity model (32GB)
M6, base16GBYesNo
M6, mid-tier24GBYesNo, under the reported requirement
M6, configured up32GBYesYes
M5 Pro, base24GBYesNo, under the reported requirement
M5 Pro, configured up48GB or 64GBYesYes, with headroom for other work

Apple silicon was missing from the last announcement

The Mac angle matters because Portable Computer, announced with Nvidia on 25 August, explicitly excluded Apple silicon. It needs an Nvidia RTX card with at least 24GB of VRAM or a DGX Spark, and Perplexity said Apple hardware was not on the roadmap. We covered that exclusion in Perplexity Partners With Nvidia to Launch Portable Computer. Hybrid mode is the first sign that Mac owners get a local-inference story of their own, even if it is a partial one.

Privacy Gate: How Hybrid Mode Handles Personal Data

perplexity hybrid mode mac local models subtasks e solid bucket curved handle

The second component in the leak is Privacy Gate, a separate local model whose only job is to inspect data before it is sent to the cloud. TestingCatalog’s description is short: if Privacy Gate “detects personal or sensitive information, Computer can ask the user how to handle it before transmission”. Three things about that design deserve attention.

A second model, not a rule set

Privacy Gate is described as a model, not a pattern matcher. That is a meaningful choice. Regular expressions catch National Insurance numbers and card numbers; they do not catch a paragraph that describes a named employee’s health condition in plain prose. A model can, because natural language processing reads meaning rather than patterns, at the cost of being probabilistic, which means it will sometimes miss and sometimes over-flag. Hybrid mode therefore ships with a privacy control that is more capable and less predictable than a classic data-loss-prevention filter.

Ask, not block

The reported behaviour is that Computer asks the user how to handle flagged data rather than refusing to send it. That fits the connector approvals Perplexity added on 24 August, where an “Always ask” setting lets users pause actions before execution. It also means the protection is only as good as the person clicking through the prompt, which any IT team that has watched users dismiss warnings will recognise. Hybrid mode gives the user a decision point; it does not take the decision for them.

How it differs from the June design

At Computex on 2 June 2026, Perplexity demonstrated what it called hybrid agentic inference, where a compact model on the device decides which parts of a task stay local and which go to the cloud, with sensitive data such as financial records and health information kept on the device by design. That system was demonstrated on Intel Core Ultra Series 3 processors and was described as exclusive to the Windows app, arriving in July. Hybrid mode on Mac appears to split the same idea into two models, one for delegation and one for inspection, which is a cleaner separation of concerns.

Design elementHybrid agentic inference (Computex, June)Hybrid mode on Mac (leak, August)
Who decides where work runsCompact on-device router modelCloud orchestrator delegates suitable subtasks
Privacy controlRouter keeps sensitive data local by designSeparate Privacy Gate model inspects outbound data and asks
PlatformWindows app, Intel Core Ultra Series 3 demoMac app, Apple silicon
Local models namedNot namedPerplexity model, Qwen 32B, Gemma-based
Stated goalBalance intelligence, accuracy, privacy and costSave costs; screen PII before transmission
StatusAnnounced for July 2026Hidden in the app, no date

The Credit Economics Hybrid Mode Is Built to Fix

perplexity hybrid mode mac local models subtasks f solid paper plane

TestingCatalog’s own reaction to the June announcement was that a local and cloud split “would be huge” if it drove Computer costs down, “since it is one of the top blockers for many at this moment”. The credit meter is the context for everything in this story, so it is worth setting out the numbers before judging how much Hybrid mode can change them.

What the meter charges

A Max subscription costs $200 a month and includes 10,000 credits, which works out at two cents per credit. Auto-refill defaults to a further $200 when credits run out, configurable up to a $2,000 cap. Published burn figures give a sense of what those credits buy: a DataCamp reviewer’s single research task that touched eight tools over just under eight minutes consumed 225.71 credits, and a complex due-diligence workflow burns around 500. At two cents a credit those are $4.51 and $10.00 respectively, and a 10,000-credit month buys roughly twenty of the heavier runs.

What one job costs under each path

The chart below puts those figures against the only other local-inference price Perplexity has published, the roughly $0.415 per task it quoted for Portable Computer escalating hard steps to Claude Opus 5. Hybrid mode has no published per-task figure yet, which is why it does not appear as a bar; the point of the chart is the size of the gap that any local delegation has to work on.

Cost of one job at the published rates ($10.00 = 100%)
500-credit due-diligence workflow, cloud Computer $10.00
225.71-credit research task, cloud Computer $4.51
Portable Computer task with cloud adviser escalation $0.415

Four cost levers in six weeks

Hybrid mode is the fourth cost lever Perplexity has been seen working on since July, and reading them together shows a company attacking the same complaint from every side. Our AI token cost calculator is a useful companion for modelling what a local worker saves on your own task mix.

LeverSeenHow it cuts the billWho it reaches
OpenRouter connection21 Jul 2026, in testingRoutes model calls through the user’s own OpenRouter key instead of creditsAnyone with an API budget
Effort selector23 Aug 2026, in testingLets the user pick a lower effort level so a task consumes fewer creditsAll Computer users, tier unknown
Portable Computer25 Aug 2026, liveWhole agent runs locally; local work consumes no creditsNvidia 24GB GPU or DGX Spark owners
Hybrid mode31 Aug 2026, hiddenCloud agent delegates suitable subtasks to a local Mac modelMacs with 16GB or 32GB RAM

The effort selector is the closest relative, and we examined its credit arithmetic in Perplexity Is Testing a Granular Effort Selector for Perplexity Computer. The selector reduces how hard the cloud thinks; Hybrid mode reduces how often the cloud is asked at all.

Hybrid Mode Against Portable Computer

The obvious comparison is with the product Perplexity launched six days earlier. Both put a model on the user’s hardware, but they answer different questions, and the difference is the clearest guide to who Hybrid mode is for.

Where the orchestrator lives

Portable Computer moves the entire stack to the local GPU: model, inference engine, agent harness, tools, connectors and sandbox. The cloud is an optional adviser that the local harness escalates to when a step is too hard. Hybrid mode keeps the orchestrator in the cloud and treats the Mac as a worker. The practical consequence is that a Portable Computer task can complete with the network unplugged, and a Hybrid mode task cannot.

What each one saves

Portable Computer’s local work consumes no credits at all, because nothing touches Perplexity’s meter unless the harness escalates. Hybrid mode saves whatever share of subtasks the orchestrator judges suitable for the local model, and that share is unknown until the feature ships. On a 16GB Mac running the 5.6GB Gemma-based model, it is reasonable to expect that share to be small, because the orchestrator will not send difficult reasoning to a worker of that size.

What each one demands

Portable Computer wants a 24GB Nvidia card, which rules out every Mac and most laptops. Hybrid mode wants a Mac with 16GB, or 32GB for the stronger models, which describes a large fraction of the machines already on desks. That trade, less saving in exchange for far broader hardware, is the whole design.

FactorPortable ComputerHybrid mode on Mac
OrchestratorLocalCloud
Local modelQwen 3.8 27B, PPLX 27B, Nemotron 3.5 Lightning soonPerplexity model, Qwen 32B, Gemma-based
Hardware floorNvidia RTX 24GB VRAM or DGX SparkMac with 16GB RAM, 32GB for the larger models
Offline capableYesNo
Credit savingAll local work is freeOnly delegated subtasks, share not stated
Privacy modelEverything starts local; asks before cloudPrivacy Gate inspects outbound data and asks
AvailabilityLive on Linux, Windows in SeptemberHidden, undated

How Hybrid Mode Fits the Cloud Runtime

Hybrid mode does not replace the infrastructure Perplexity has built for Computer this year; it plugs into it. Two pieces of that infrastructure explain why delegating subtasks is straightforward for Perplexity to add now.

SPACE separates the session from the sandbox

On 15 July 2026 Perplexity described SPACE, its Sandboxed Platform for Agentic Code Execution, which it said powers every Computer session. Each task runs in a disposable Firecracker microVM with its own guest kernel, and rolling snapshots preserve memory and files so a session can pause, resume or branch. The stated design principle was that SPACE “separates the session from the sandbox running it”. A runtime built to move a session between sandboxes is a runtime that can, in principle, treat a Mac as one more place a subtask executes.

The runtime got faster before it got wider

Perplexity’s published SPACE figures show median sandbox creation falling from 185 milliseconds to 60, and 90th-percentile latency from 447 milliseconds to 89. Those improvements matter for Hybrid mode because a subtask delegated to a local model still has to be handed off from a cloud session and its result handed back; a slow runtime would eat the saving. The chart scales all four figures against the slowest.

SPACE sandbox latency, before and after (447ms = 100%)
90th percentile, before 447ms
Median, before 185ms
90th percentile, after 89ms
Median, after 60ms

Memory and subagents already exist

Two other June and August additions complete the picture. The Brain memory system, released on 19 June, gives Computer persistent memory across sessions, and the 24 August release notes made GPT-5.6 Terra the default subagent model. Hybrid mode slots a local worker into a system that already has memory, subagents and a portable sandbox; the novelty is the location of the worker, not the architecture around it.

What Hybrid Mode Means for UK Businesses

For a UK business that has been watching Perplexity Computer with interest and its credit bills with alarm, Hybrid mode is worth planning around but not worth waiting for. Three practical points follow from the leak.

Privacy Gate is a control, not a compliance answer

A local model that flags personal data before it leaves the machine is a genuinely useful control, and it is better than the nothing most agent products offer. It is not a lawful basis, a data processing agreement or a transfer risk assessment, and it does not change where the cloud side of Perplexity Computer processes data. Treat Privacy Gate as one layer in a wider approach to agent governance, of the kind we set out in AI agents that pass authentication can still drift, expose data or get memory-poisoned.

Buy the 32GB Mac mini, not the 16GB one

If a pilot is on the cards, the reported requirements make the purchasing decision simple. The 16GB machine limits Hybrid mode to the 5.6GB Gemma-based model, which will take the fewest subtasks. A 32GB configuration unlocks both stronger local models and, because Perplexity recommends a Mac mini for continuous operation anyway, an always-on desktop is the sensible shape for the pilot. Model the saving against your actual burn profile rather than the headline, and cap auto-refill while you measure.

Fit it into an AI strategy, not a procurement line

Delegating subtasks to local hardware is a pattern that will outlast this product. Apple, Microsoft and Google are all pushing on-device models, and Perplexity’s own CEO argued at Computex that “you don’t want all your compute centralized in servers”. A business that decides now where its sensitive data may be processed, and which tasks justify frontier-model cost, will be ready for Hybrid mode and for whatever ships after it. That is a decision for an AI strategy, and it applies equally to the autonomous AI agents a business builds for itself.

The Limits of the Hybrid Mode Story

Everything above rests on a report about a hidden feature, and the caveats are as important as the details.

It is unreleased and undated

TestingCatalog is explicit that the features “remain hidden” and that Perplexity has announced neither a release date nor whether access will depend on subscription tier. The June hybrid agentic inference feature was promised for July and, on the Windows side, was still being described as arriving weeks later. Hybrid mode on Mac could ship next week or slip into the autumn, and a business plan should assume the latter.

Small models are still small models

The strongest evidence about what a local worker of this size can do comes from Perplexity’s own Portable Computer benchmarks, where the local Qwen configuration scored 59.6% on Terminal Bench 2.1 and reached 73.0% only when hard steps were escalated to Claude Opus 5. A 32B model on a Mac will sit in the same range, and the 5.6GB option lower. Hybrid mode will save money precisely to the extent that the orchestrator is disciplined about what it delegates.

The saving is unquantified

“Save costs” is the only claim Perplexity’s interface makes, and there is no per-task figure, no credit rate for local work and no statement of which subtasks qualify. Until the feature ships and users publish burn figures, the honest position is that Hybrid mode will reduce credit consumption by an amount nobody outside Perplexity can yet estimate.

Hybrid Mode FAQs

What is Hybrid mode in Perplexity Computer?

Hybrid mode is an unreleased setting in the Perplexity Computer Mac app, surfaced by TestingCatalog on 31 August 2026, that lets the cloud orchestrator delegate suitable subtasks to a model running locally on the Mac in order to save credits.

Which local models does Hybrid mode offer?

Three downloadable options were found: a proprietary Perplexity model of about 19GB, a Qwen 32B model of about 17.4GB, both needing 32GB of RAM, and a smaller Gemma-based model of about 5.6GB that runs on Macs with 16GB.

What is Privacy Gate?

Privacy Gate is a separate local model that inspects data before it is sent to the cloud. If it detects personal or sensitive information, Computer can ask the user how to handle it before transmission.

Does Hybrid mode work offline?

No. Orchestration stays in the cloud, so Hybrid mode needs a connection. Portable Computer, which runs the whole agent on an Nvidia GPU, is the offline-capable product, and it does not support Apple silicon.

Which Mac should I buy for Hybrid mode?

A 32GB configuration unlocks all three local models; a 16GB Mac limits you to the Gemma-based option. Perplexity recommends a Mac mini because it lets the agent run continuously.

When will Hybrid mode be released?

Perplexity has not announced a date or said which subscription tiers will get it. The feature is hidden in the current app.

References and Further Reading