Imagine Agent is the part of Grok that stopped asking you for one prompt at a time. xAI put it into beta at the end of April 2026 as an open canvas with an agent sidebar, where you describe a project — a product photoshoot across several SKUs, a one-minute short film, a brand identity set — and the agent plans it, generates the pieces, edits them in place and assembles the result. What it could not do at launch was draw better than the image model sitting underneath it. That model has now changed.

On 7 August 2026 xAI shipped Imagine Image 2.0 as the new Quality Mode on grok.com/imagine and inside the Grok iOS and Android apps. It is a different class of image model from the one the canvas started on: it plans typography and layout the way a designer would, keeps small text legible inside dense compositions, and preserves elements you supply across every subsequent generation and edit. Because Quality Mode is the tier the canvas draws from, the Imagine Agent inherits all of it — precision region editing, background removal, five reference images per generation and Smart Resize across nine aspect ratios.

We covered the model’s consumer-facing side when it landed, in our write-up of Grok Imagine’s photo editing and resizing update. This article is about the other half: what a better image model does to a workspace built for autonomous AI agents, what the Arena scores really say, what the API now costs per image, and which of the agent’s original limits are still real in September 2026.

What the Imagine Agent Is and Why Its Model Swap Matters

imagine agent grok image 2 0 model update b camera block with round lens barrel

The Imagine Agent is not a new model. It is an orchestration layer — a canvas, a planner and a set of in-place editing tools — that calls whatever image and video models Grok Imagine exposes. That distinction is the whole reason a model upgrade is news rather than a footnote.

The canvas that replaced the prompt box

TestingCatalog reported the Imagine Agent on 30 April 2026 as a gradual beta rollout to Grok Heavy and Super Grok subscribers who already had Imagine privileges. Elon Musk announced Agent Mode publicly on X the following day, and the feature reached the web app over the days after that.

Instead of a chat thread, you get an open canvas. Third-party walkthroughs describe a node-based workflow where each step is an object you can click and re-prompt without restarting the pipeline, branch for A/B variations, or feed with reference images dropped onto the canvas. The toolset around it is deliberately editor-like: Quick Animate to turn a still into a short clip, Group Images to treat several shots as one set, Crop, Trim and Fade for in-canvas edits, a Select and Hand tool for arranging, and Stitch to chain clips into longer sequences.

From single prompts to delegated projects

The behaviour that makes it an Imagine Agent rather than a fancy gallery is delegation. Ask for a one-minute short and it drafts a scenario, generates the individual scene clips, stitches them into a sequence and produces a companion poster image. Ask for a product story and it runs the shoot across your SKUs. The published preset workflows — Create Worlds, Short Film, UGC Product Stories and Brand Identity — are packaged versions of exactly that loop.

Why the underlying image model decides the ceiling

An agent that plans ten steps and renders each one badly is worse than a single good prompt, because it multiplies the error. Every weakness in the base model compounds across a multi-step project: text that smears at small sizes ruins a poster, a subject that drifts between frames ruins a storyboard, and a logo that mutates ruins a brand set. This is precisely the surface Image 2.0 targets, which is why the model swap changes what the Imagine Agent can be trusted with rather than just how its outputs look.

DateWhat shippedEffect on the Imagine Agent
30 Apr 2026Imagine Agent beta on an open canvasMulti-step projects arrive, on the then-current Quality Mode model
1 Aug 2026Imagine Video 1.5 adds character references and 1080pRaises the video ceiling the agent stitches against
7 Aug 2026Imagine Image 2.0 becomes Quality ModeThe agent’s image tier gains typography, layout and precision editing
28 Aug 2026API adds auto quality, five source images, 21:9 and 5:2The same capability becomes scriptable outside the canvas
2 Nov 2026grok-imagine-image-quality retiresOld slug is served by Image 2.0 at low quality, cheaper per image

Imagine Image 2.0: The Model Now Behind the Imagine Agent

imagine agent grok image 2 0 model update c funnel cone with straight spout

xAI’s own announcement is unusually specific about what the model was built for, and it is not “prettier pictures”. The pitch is instruction fidelity and design-grade control — the things a delegated workflow needs and a one-off prompt can survive without.

Typography and layout that survive small sizes

The company’s claim is that Image 2.0 “plans typography and layout the way a designer would”, renders dense multi-part visuals with small text still sharp, and preserves supplied elements across new generations and edits. For an Imagine Agent building a brand set, that last clause is the load-bearing one: a logo, a product shot or a character has to come back unchanged in step nine having been introduced in step two.

The editing toolset the canvas inherits

Editing in Image 2.0 is a first-class feature rather than a bolt-on to a text-to-image model, and every tool below is available to the Imagine Agent working on the same canvas. Segmentation and background removal are computer vision problems dressed up as buttons, and they are the two that most often decide whether a generated asset can be dropped into an existing template or has to be redone by hand.

ToolWhat it doesWhy it matters to an agent workflow
Magic WandEdits the region you point at, leaving the rest untouchedFixes one step without re-rolling the whole project
SegmentationSelects a precise area for modificationMakes automated edits repeatable across a set
Background removalExports the subject on transparencyProduces compositable assets, not flat pictures
Multi-reference editingUp to five input images in one generationCarries brand, character and product identity forward
Smart ResizeRecomposes one image into other aspect ratiosOne approved asset becomes a full channel set

Smart Resize and the nine ratios

Smart Resize recomposes rather than crops, which is the difference between a usable banner and a decapitated one. xAI lists nine ratios in the product: 1:2, 9:16, 2:3, 3:4, 1:1, 4:3, 3:2, 16:9 and 2:1. The API has since added two more, 21:9 for cinematic widescreen and 5:2 for wide banners, taking the scriptable total to eleven. Sixteen workflow templates ship alongside, covering photo edits, product changes, e-commerce, headshots, icons, sprites, emojis, game assets and merchandise.

Imagine Agent Benchmarks: What the Arena Scores Actually Show

imagine agent grok image 2 0 model update d set square triangle standing upright

The headline is that Image 2.0 ranks second in the world in both text-to-image generation and image editing. The more useful number is how far it moved from the model the Imagine Agent used to run on.

The jump from the old Quality Mode model

On the Arena leaderboards as of 8 August 2026, Image 2.0 scored 1320 in text-to-image and 1439 in image editing. The outgoing grok-imagine-image-quality model scored 1228 and 1390. OpenAI’s gpt-image-2 held first place on both boards with 1380 and 1463.

Arena points gained versus points still missing (from the figures above)
Text-to-image gain over the old model, 1228 to 1320 +92
Text-to-image gap still behind gpt-image-2, 1320 to 1380 60
Image-editing gain over the old model, 1390 to 1439 +49
Image-editing gap still behind gpt-image-2, 1439 to 1463 24

Read that chart the way an operator would rather than the way a leaderboard does. The editing board is where the Imagine Agent spends most of its time, and it is also where the remaining gap is smallest — 24 points, against a 49-point gain. On generation the model moved further but still trails by 60. In other words, the upgrade closed more ground on the capability the canvas actually leans on.

Reading the leaderboard honestly

Arena scores are human preference ratings collected from side-by-side comparisons of AI models, not a measure of whether an asset is usable in your brand system. Second place with a 60-point deficit is not the same as parity, and a preference score says nothing about how the model behaves over a ten-step Imagine Agent project where consistency matters more than any single frame. Treat the numbers as evidence the gap narrowed, then test the model on your own repeat-element workload before deciding anything.

What Changes Inside the Imagine Agent Canvas

imagine agent grok image 2 0 model update e dial knob on round base plate

Most of what the model swap delivers to the canvas is not a new button. It is the same buttons producing results that survive being chained together.

Reference images went from three to five

The consumer apps accept up to five input images in a single generation, and on 28 August 2026 the API’s editing endpoint was raised to match — up to five source images per request, previously three. That is a 67% increase in the identity you can pin per call, and for an Imagine Agent doing brand work it is the difference between carrying a logo and carrying a logo, a product, a colour reference, a model and a set.

Templates as packaged agent workflows

The sixteen Image 2.0 templates and the agent’s own presets — Create Worlds, Short Film, UGC Product Stories, Brand Identity — do the same job from opposite ends. One packages a single well-formed edit; the other packages a project shape. Together they mean a first useful output rarely requires prompt engineering, which matters more for a delegated agent than for a person iterating in a chat window.

What the Imagine Agent still cannot do

Third-party testing from the May 2026 rollout recorded real limits, and it is worth being precise about which ones the August upgrades have already overtaken.

Limit reported at launchStatus in September 2026
Video clips capped at 6 seconds, 720pPredates Imagine Video 1.5, which added native 1080p and character references on 1 August
Character drift across stitched clipsPartly addressed by character references and five-image conditioning; still worth testing
No audio generation in the agent flowReported at launch; Imagine Video 1.5 is documented with native audio, so verify in your own account
Web-only, no iOS or Android equivalentStill web-only; Image 2.0 Quality Mode itself is on all three platforms
Content policy blocks real identifiable people and protected IPUnchanged, and the correct behaviour for commercial use

The honest summary: the Imagine Agent’s image ceiling has moved decisively, its video ceiling has moved, and its platform ceiling has not moved at all.

Imagine Agent Access, Tiers and Platform Limits

imagine agent grok image 2 0 model update f anvil block with tapered horn

Access is where a lot of enthusiasm meets a paywall. xAI has not published a tier page specific to the Imagine Agent, so the figures below come from TestingCatalog’s rollout reporting and third-party guides written during the beta — treat them as a guide to check against your own billing page, not as a price list.

PlanReported priceImagine Agent access
Free Grok£0No canvas; Imagine only
SuperGrok Lite$10 / monthNot included
SuperGrok$30 / month or $300 / yearFull access, capped video renders
SuperGrok Heavy$300 / monthHighest video caps, priority access

Turning the canvas on

Reports from the beta describe two routes into the Imagine Agent: the dedicated grok.com/imagine/agent path, or an Agent Mode toggle in the Imagine input field. Both land you in the same canvas with the agent sidebar attached.

Web-only, and what that costs a mobile team

Image 2.0 Quality Mode runs on web, iOS and Android. The Imagine Agent does not — it is a browser feature. For a marketing team that shoots and approves on phones, that is a genuine workflow split: generate and edit anywhere, but orchestrate a project only at a desk. Anyone planning around the agent should assume desktop for the foreseeable future and design the approval step accordingly.

The Imagine Agent API: Model Slugs, Quality Tiers and Costs

When Image 2.0 launched, the API was described as planned but not yet available, and the existing Imagine API kept serving the older quality-tier model. That gap has since closed, and the pricing is now the most concrete thing published about the model.

Model slugPrice per imageStatus
grok-imagine-image-2.0$0.04Current, powers Quality Mode
grok-imagine-image-quality$0.05Retires 2 November 2026
grok-imagine-image$0.02Version 1.0, unaffected by the retirement
Published list price per image, relative to the outgoing quality model
grok-imagine-image-quality, retiring 2 November $0.05
grok-imagine-image-2.0, the model behind the canvas $0.04
grok-imagine-image, version 1.0 $0.02

The better model is a penny cheaper per image than the one it replaces. On a 20,000-image month that is a $200 difference in the right direction — small in isolation, but notable because upgrades in this market usually move the other way.

The auto quality parameter, and the billing catch

On 28 August 2026 the API’s quality parameter on grok-imagine-image-2.0 gained an auto setting, and auto became the default when the parameter is omitted, replacing medium. Auto currently resolves to low for image generation and medium for image editing. The line to read twice is that images are billed at the quality they are served at — so a request that omits the parameter is now priced by whatever auto picked, not by the tier you assumed. You can still pass low or medium explicitly, and on any automated Imagine Agent-style pipeline you should.

The November retirement, and what it quietly does to price

From 2 November 2026 the grok-imagine-image-quality slug is discontinued. Requests to it will be served by grok-imagine-image-2.0 with quality set to low, with no change to the request or response structure and a reduced per-image price. Version 1.0 is unaffected. That is a rare shape of deprecation: existing integrations keep working, get a better model, and pay less — provided low quality is genuinely acceptable for the workload, which for thumbnail and variant generation it often is and for final assets it often is not.

Per-image economics worth modelling

Two figures decide whether an API-driven pipeline is cheaper than a canvas subscription. At $0.04 per image, a SuperGrok seat at $30 a month is worth roughly 750 images before the API becomes the cheaper unit — and the seat also buys the Imagine Agent canvas, the video tools and a human in the loop. The API wins on volume and on integration; the canvas wins on judgement. Most teams that adopt the Imagine Agent seriously end up paying for both, which is worth putting in the budget from the start rather than discovering in month three.

What the Imagine Agent Means for Business Creative Work

Strip away the model numbers and the practical claim is narrow: a category of small, repetitive, previously human production work is now automatable to a standard that survives review.

Where it genuinely saves money

Multi-SKU product imagery, channel resizing, storyboard and concept passes, headshot and icon sets, and first-draft campaign variants are the clear wins. All five share a shape — many outputs, one identity to preserve, a human approving at the end. That is exactly what five-image conditioning plus Smart Resize plus a planning agent is built to do, and it is close to what a workflow automation programme would try to automate in any other function.

The governance questions to settle first

Before an Imagine Agent touches production creative, settle four things. Who signs off generated assets, given the agent produces dozens per run. Where the licensing sits for output used commercially, and what your contract with xAI actually says about it. How you keep prompts and reference images containing unreleased product out of a third-party service. And what your disclosure policy is when an asset in a customer-facing campaign is synthetic.

None of that is a reason to avoid the tool. It is the same list any organisation works through when it hands a repeatable task to software, and it belongs in a written AI strategy rather than in a campaign brief. The answers are much easier to draft before the first campaign than after it.

A sensible pilot

Pick one recurring, low-risk asset family — social variants of an approved hero image is the usual candidate. Run it through the Imagine Agent for a month against a measured baseline, with the same discipline any intelligent automation pilot needs: hours spent, revisions requested, assets rejected at review. If precision editing and reference conditioning hold identity as well as the marketing claims, the numbers will show it inside four weeks. If they do not, you have spent one seat finding out.

Imagine Agent FAQ: Access, Models, Video and API

Is the Imagine Agent the same thing as Grok Imagine?

No. Grok Imagine is the image and video generation product; the Imagine Agent is a beta canvas mode inside it that plans and executes multi-step projects. Imagine works everywhere Grok does. The agent canvas is web-only.

Which model does the Imagine Agent use now?

It draws on Grok Imagine’s Quality Mode, which since 7 August 2026 is Imagine Image 2.0. xAI has not published a separate announcement tying the agent to the model, so treat the connection as following from Quality Mode rather than from a dedicated release note.

Is Image 2.0 available in the API?

Yes. The slug is grok-imagine-image-2.0, listed at $0.04 per image, with a quality parameter accepting auto, low and medium. At launch on 7 August the API was described as planned but not yet shipped, so older coverage saying otherwise is simply out of date.

What happens to grok-imagine-image-quality?

It retires on 2 November 2026. Requests keep working and are served by grok-imagine-image-2.0 at quality low, with the same request and response structure and a lower per-image price.

Can the Imagine Agent produce video with sound?

Grok Imagine Video 1.5 is documented with native audio and, since 1 August 2026, native 1080p and character references. Third-party testing of the agent canvas during its May beta reported no audio and a 6-second, 720p ceiling. Those reports predate the video upgrade, so verify the current behaviour in your own account before planning around either answer.

How many reference images can it hold?

Five, in both the consumer apps and, since 28 August 2026, the API’s editing endpoint — up from three on the API side.

Is the output safe to use commercially?

The content policy blocks real identifiable people without consent, protected IP and illegal imagery. Beyond that, commercial licensing is a contractual question between you and xAI, and it should be confirmed in writing before an Imagine Agent output appears in a paid campaign.

References