Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking are the two model names that began appearing on Google Cloud’s console quotas page on the morning of 15 September 2026, according to TestingCatalog, which posted screenshots on Threads and credited the find to researcher Bedros Pamboukian. Neither model is public. Google has not announced either name, and its developer release notes still end with the Gemini 3.8 Flash and Lyria 3.5 launches of 2 and 3 September.

The same morning, TestingCatalog folded the find into a three-line “model forecast”: Gemini 3.8 Live flagged “Today?”, Grok 4.7 flagged “This week?” and expected to land “near Opus 5.0, not 5.1”, and a tester report that some Claude Code traffic is reaching an unofficial Opus 5.2. For teams building voice products and AI agents, the first line matters most, because the newest conversational model Google offers developers is still a March preview.

This article sets out what the quota-page sighting does and does not prove, where Gemini 3.8 Live would sit in Google’s Live API lineup, how it compares with OpenAI’s GPT-Live-1, and what Elon Musk’s posts and three prediction markets say about Grok 4.7. Every price, date and probability below was read on 15 September 2026, and anything that is inference rather than published fact is labelled as such.

What TestingCatalog Found: Gemini 3.8 Live on the Cloud Quota Page

Gemini 3.8 Live - gemini 3 8 live slugs cloud quota page grok 4 7 forecasts b tall object hidden under a draped cloth

The report is short, so it is worth reading exactly. TestingCatalog wrote that “Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking model names have started to appear on the GCP Console quotas page”, and added that “a new Gemini Live model, based on the latest Gemini 3.8 Flash, is expected to be a big leap”. The post carried two screenshots of the console list and no link to a Google source.

The two posts behind the headline

The first post is the sighting itself. A second post in the same thread supplies the context: “So far, Gemini 3.1 Live Preview is the latest Gemini Live model available via APIs.” It notes that Google recently added connector calls and Deep Research to Gemini Live, and that “OpenAI also released the GPT Live 1 model on the APIs last week, and it seems like Google has a response to that”. The credit line reads “Discovered by Bedros Pamboukian”.

A little later, TestingCatalog’s daily “MODEL FORECAST” post compressed the finding into one line: “Today? Gemini 3.8 Live and Live Extended Thinking slugs appeared on the Cloud quota page; not public.” That is where the word “slugs” in most headlines comes from. The question mark matters. It is a guess about timing, not a report that Gemini 3.8 Live has launched.

What the sighting establishes, claim by claim

Each part of the Gemini 3.8 Live report carries a different level of evidence. Here is how the claims stood at 12:00 UTC on 15 September.

ClaimSourceStatus
“Gemini 3.8 Live” and “Gemini 3.8 Live Extended Thinking” appear on a GCP Console quotas pageTestingCatalog screenshotsReported with screenshots; not independently viewable
The models are not publicTestingCatalog forecast postConsistent with Google’s docs: no model page, price or release note
Gemini 3.8 Live is based on Gemini 3.8 FlashTestingCatalogInference from the name; Google has not said
“Expected to be a big leap”TestingCatalogOpinion; no benchmark published
A response to OpenAI’s GPT-Live-1TestingCatalogFraming; the timing fits, the motive is unconfirmed
Launch “Today?”TestingCatalog forecast postNo launch on any Google page by midday UTC

What we could not verify

We could not open a Google Cloud project that shows the entries, so we cannot confirm the exact model identifiers, which quota metrics they sit under, or whether every project sees them. TestingCatalog’s wording moves between “model names” and “slugs”, and the screenshots show display names.

The identifier Google eventually ships for Gemini 3.8 Live may look different. Today’s Live models are called gemini-3.1-flash-live-preview in the Gemini API and gemini-live-2.5-flash-native-audio on Google Cloud, two quite different naming patterns, and neither puts “Live” straight after the version number the way the console entries do.

What a Quota Page Entry Proves About Gemini 3.8 Live

gemini 3 8 live slugs cloud quota page grok 4 7 forecasts c water tower tank on four straight legs

A quota is a limit on how much of a resource a Google Cloud project may use. Google’s quota documentation says quotas “generally apply at the Google Cloud project level”, and its model tables track limits per base model through a base_model dimension. The published table lists requests for gemini-embedding, for example, under the metric global_embed_content_input_tokens_per_minute_per_base_model.

Why model names surface on quota pages early

Before any customer can call a model at scale, Google has to create the metrics that meter it. Those entries are configuration, and configuration tends to ship before the launch blog post, the model card, the pricing row and the documentation page. That makes the quotas page a useful early signal, and it is why watchers such as TestingCatalog check it.

It is also why an entry proves less than it seems. It shows that serving infrastructure for Gemini 3.8 Live has been set up. It does not show that access has been granted, that pricing is final or that a launch date has been set. Quota entries for a model can sit unused for weeks, and a model can be renamed between configuration and launch.

A ladder of release signals

Release evidence comes in steps, from internal configuration to a public announcement. Gemini 3.8 Live has reached only the second rung.

SignalWhat it showsSeen for Gemini 3.8 Live?
Internal testing reportsA model exists and is being evaluatedImplied by the name, not reported
Quota page entryServing limits have been configuredYes, per TestingCatalog
Model returned by the API model listSome keys can call itNot reported
Pricing page rowBilling is definedNo, checked 15 September
Model documentation pageGoogle supports it publiclyNo
Release notes entryLaunchNo; newest entries are 2 and 3 September
Launch blog postAnnouncement with claims and benchmarksNo

How quickly the last leak turned into a launch

For Gemini 3.8 Flash, the gap between the first public report of testing and general availability was short. Business Insider reported on 27 August that Google employees were already testing the model, and Google made gemini-3.8-flash generally available on 2 September, six days later.

That is one data point, not a rule, and a quota entry is a different kind of signal from staff testing. A Live model also has more to prove than a text model before launch: latency, turn-taking, interruptions and audio quality all have to hold up in real conversations, not just in a benchmark harness. Gemini 3.8 Live could follow in days or sit in configuration for weeks.

Where Gemini 3.8 Live Would Fit in Google's Live API Lineup

gemini 3 8 live slugs cloud quota page grok 4 7 forecasts d speech bubble block with one pointed tail

The Live API is Google’s interface for real-time conversation. Google describes it as a stateful connection over WebSockets that takes in a continuous stream of audio, video frames or text and replies with audio or text. The models behind it have their own names and release cadence, separate from the Flash and Pro text models.

Under the hood, these models fold speech recognition, natural language processing and speech generation into one model instead of a pipeline of three, which is where most of the latency saving comes from. Gemini 3.8 Live would be the first Live model to carry the 3.8 version number.

The Live models Google lists today

Google’s two developer platforms list different Live models, which is part of why a new name is notable. This is the lineup Gemini 3.8 Live would join.

ModelPlatformStatusWhat it does
gemini-3.1-flash-live-previewGemini APIPreview, released 26 March 2026Audio-to-audio dialogue; text, image, audio and video in; text and audio out
gemini-live-2.5-flash-native-audioGoogle CloudGenerally available, marked “Recommended”Low-latency voice agents with affective dialog, proactive audio and tool use
gemini-2.5-flash-native-audio-preview-12-2025Gemini APIPreview, released 12 December 2025Earlier native audio model for the Live API
gemini-3.5-transcribe-live-previewGoogle CloudPreviewSpeech-to-text only; returns text and does not hold a conversation
gemini-3.5-live-translate-previewGemini APIPreviewReal-time speech-to-speech translation in 70+ languages
Gemini 3.8 Live and Gemini 3.8 Live Extended ThinkingConsole quotas page onlyNot publicUnknown

What Gemini 3.1 Flash Live can and cannot do

The current preview sets the bar that Gemini 3.8 Live has to clear. Its model page lists an input limit of 131,072 tokens and an output limit of 65,536, with function calling, Google Search grounding and thinking supported. Structured outputs, code execution and caching are not supported.

The page also lists gaps. Function calling is synchronous only, so the model “will not start responding until you’ve sent the tool response”. And “proactive audio and affective dialogue” are “not yet supported in Gemini 3.1 Flash Live”. The older 2.5 native audio model on Google Cloud does support both, so developers currently choose between a newer reasoning core and richer conversational behaviour. A Gemini 3.8 Live that closed that gap would matter more than any single benchmark.

What Google charges for Live today

Google has published no price for Gemini 3.8 Live. These are the paid-tier list prices for the Live models it does sell, per million tokens in US dollars, with Google’s own per-minute equivalents in brackets.

ModelAudio inputAudio outputText inputText output
Gemini 3.1 Flash Live Preview$3.00 ($0.005/min)$12.00 ($0.018/min)$0.75$4.50
Gemini 2.5 Flash Native Audio$3.00$12.00$0.50$2.00
Gemini 3.5 Live Translate$3.50 ($0.0053/min)$21.00 ($0.0315/min)Not listedNot listed
Gemini 3.8 LiveNot publishedNot publishedNot publishedNot published

Google notes that output prices include thinking tokens. If an Extended Thinking variant of Gemini 3.8 Live spends more of each turn reasoning, the audio rate could stay the same while the cost of each turn rises.

Why Gemini 3.8 Live Needs Its Own Model, and What "Extended Thinking" Could Mean

gemini 3 8 live slugs cloud quota page grok 4 7 forecasts e casserole pot with closed lid and side tabs

There is a plain reason Google cannot simply point the Live API at the model it launched on 2 September. The Gemini 3.8 Flash page on Google Cloud lists “Gemini Live API: Not supported”. The text model takes text, image, audio and video in, with a 1,048,576-token context window and up to 65,536 output tokens, but it answers in text. A live voice model needs audio output, streaming turn detection and interruption handling, which is separate training and serving work. That is the gap Gemini 3.8 Live would fill.

Gemini 3.8 Flash has no minimal thinking level

The more telling clue is how thinking works. Gemini 3.1 Flash Live uses thinking levels of minimal, low, medium and high, and Google says “the default is minimal to optimize for lowest latency”.

Gemini 3.8 Flash is different. Google’s documentation says thinking_level MINIMAL “is not available for 3.8 Flash” and that setting it “will return an API validation error”. The supported values are low, medium and high. If Gemini 3.8 Live inherits that floor, even its fastest setting would think more than today’s Live default, and a voice model that thinks longer answers later.

Reading the “Extended Thinking” name

This part is inference, because Google has said nothing. The plainest reading is that Gemini 3.8 Live is tuned for conversational speed, and Gemini 3.8 Live Extended Thinking allows longer reasoning per turn for tasks such as working through a support case, checking a calculation or planning a multi-step action while the caller waits. Voice products already live with that trade-off: a slow answer feels broken, but a fast wrong answer is worse.

The label is also unusual for Google. Its API talks about “thinking levels” and, before that, “thinking budgets”, while “extended thinking” is the phrase Anthropic uses for Claude. Quota-page names are internal, so the Gemini 3.8 Live Extended Thinking name may not survive to launch in that form.

Two variants, two jobs

If the two names do map to a fast and a deliberate model, developers would route calls between them in much the same way they already choose thinking levels. This is how the split could look, stated as a working hypothesis rather than a specification.

NeedLikely fitWhy
Receptionist, booking, simple FAQsGemini 3.8 LiveTurn latency matters more than depth
Technical support with diagnosisGemini 3.8 Live Extended ThinkingWrong answers cost more than a pause
Tutoring and worked problemsExtended ThinkingChecking steps is the product
Live translationNeither; Gemini 3.5 Live Translate existsA dedicated model is already on sale

Gemini 3.8 Live Versus OpenAI's GPT-Live-1

gemini 3 8 live slugs cloud quota page grok 4 7 forecasts f two dice cubes side by side with a gap

TestingCatalog framed the sighting as Google’s answer to OpenAI, and the timing supports that reading. OpenAI brought GPT-Live-1 to its API on 10 September 2026, 64 days after GPT-Live launched in ChatGPT on 8 July, and we covered that launch in detail in our GPT-Live-1 API article. Five days later, the Gemini 3.8 Live names appeared in Google’s console.

What is known on each side

The fair comparison today is three-way: OpenAI’s shipped model, Google’s shipped preview, and the unreleased Gemini 3.8 Live.

FactorGPT-Live-1Gemini 3.1 Flash Live PreviewGemini 3.8 Live
StatusGenerally available in the API since 10 September 2026Preview since 26 March 2026Names on a quotas page
Interface/v1/live/sessionsLive API over WebSocketsPresumably the Live API
List price$0.05 per minute of session time, billed per second$0.005/min audio in, $0.018/min audio outNot published
Context128,000 tokens131,072 input tokensUnknown
Harder reasoningCan hand work to a separate reasoning model; four of OpenAI’s seven launch charts used oneBuilt-in thinking levels, default minimalA separate Extended Thinking name exists
Published benchmarksYes, including tool calling and interruptionsModel cardNone

The per-minute arithmetic

The two companies price differently, so a like-for-like figure needs an assumption. Take a ten-minute call in which the app streams the caller’s microphone for all ten minutes and the model speaks for four of them.

At list rates, GPT-Live-1 costs 10 × $0.05 = $0.50, because OpenAI bills session time, silence included. Gemini 3.1 Flash Live costs 10 × $0.005 + 4 × $0.018 = $0.05 + $0.072 = $0.122, before thinking tokens, tool calls or growing context. That is about a quarter of OpenAI’s list price on this assumption, but the gap will move once Google prices Gemini 3.8 Live.

For reference, Gemini 3.8 Flash launched at $0.75 per million input tokens and $3.75 per million output tokens, the same introductory rate as 3.7 Flash, and both figures double to $1.50 and $7.50 from 1 January 2027. A newer core could well carry a higher Live rate than the March preview.

Why the list price will not decide it

A cheaper rate does not settle the choice. OpenAI published measured results for GPT-Live-1 on tool calling and interruption handling, and customer results on real call volumes. Google has published nothing for Gemini 3.8 Live, and its current preview still lacks asynchronous function calling. Until Google publishes, the comparison that matters in production is GPT-Live-1 against the models developers can actually call.

How Fast Gemini 3.8 Live Could Follow the Flash Release Cycle

Google’s Flash line has moved quickly this summer. Gemini 3.6 Flash arrived on 21 July, 3.7 Flash on 13 August and 3.8 Flash on 2 September, gaps of 23 and 20 days; we covered the latest in our Gemini 3.8 Flash in Agent Studio article. The conversational Live line has not kept pace. The newest Live model for dialogue in the Gemini API, gemini-3.1-flash-live-preview, dates from 26 March.

Counted to 15 September, the gap between the text models and the Live model is stark.

Days since each Gemini release, as of 15 September 2026
3.1 Flash Live Preview (26 March) 173 days
3.6 Flash (21 July) 56 days
3.7 Flash (13 August) 33 days
3.8 Flash (2 September) 13 days

What the cadence suggests, and what it does not

A Live model built on 3.8 Flash would close a gap of almost six months between Google’s conversational and text models. Nothing in the cadence dates the Gemini 3.8 Live launch, though. Google has not tied Live releases to Flash releases: the Live model numbered 3.1 has served while three Flash generations shipped, and the 2.5 native audio model went through dated previews in September and December 2025 before a version reached general availability on Google Cloud.

Where “Today?” stood at midday

TestingCatalog’s forecast asked whether Gemini 3.8 Live would ship on 15 September. At 12:00 UTC that day, the Gemini API release notes still showed Lyria 3.5 on 3 September as the newest entry, Google Cloud’s release notes had nothing newer on Live, and neither pricing page mentioned the model. A same-day launch was still possible, but none of the public surfaces had moved.

The Grok 4.7 Forecast: What Musk Has Promised

The second line of the forecast turns to SpaceXAI. TestingCatalog wrote: “This week? Grok 4.7 should land near Opus 5.0, not 5.1; multimodal still needs work.” That wording tracks a post Elon Musk made on 14 September. Unlike Gemini 3.8 Live, where the evidence is a console entry, the evidence for Grok 4.7 is almost entirely what one person has posted, and every date he has given has passed.

The posts that set expectations

These are the five Musk posts that set the Grok 4.7 timeline, with their view counts as read through the fxtwitter API at midday on 15 September.

Date (UTC)What Musk wroteViews
24 July“Grok 4.6 in 2 weeks and Grok 4.7 in 4 weeks”1.80 million
12 August“significantly better than 4.6 and should be ready in 3 to 4 weeks”1.33 million
2 September“Grok 4.7 comes out in 10 days”10.29 million
11 September“needs a few more days to cook”2.05 million
14 September“roughly on par with Opus 5.0, not 5.1. Better in some ways, worse in others. We need to fix multimodal performance.”1.99 million

Why the model slipped, in Musk’s words

The 11 September post is the only one that gives a technical reason. Musk wrote that xAI “might have penalized response length too much (or something) in RL, as it still gives up on hard tasks (that it can do!) too early and isn’t yet sufficiently rigorous in checking its work”.

In plain terms, reinforcement learning taught the model to keep its answers short, and it learned to stop before finishing difficult work. That kind of tuning problem can be fixed in days or weeks, which is why the forecast still says “this week”. We set out the full sequence of dates, and the 2.5-trillion-parameter Grok 4.8 Musk says comes next, in our Grok 4.8 article. Counted to 15 September, Grok 4.7 is 25 days past the four-week estimate of 24 July.

What Prediction Markets Price for Grok 4.7

Three venues let people bet on the release date. At midday on 15 September they broadly agreed that Grok 4.7 is more likely than not to ship within days, while disagreeing on how sure to be. None of them offers a contract on Gemini 3.8 Live.

Three venues, three different questions

The contracts are not identical, so the prices should not be averaged. The table below sets out what each one actually asks.

Venue and questionPrice at 12:00 UTC, 15 SeptemberActivityCatch
Polymarket: next Grok model (4.7 or higher) released by 18 September?64% (bid 63¢, ask 65¢)$15,172 tradedAny 4.7-or-higher Grok counts, including Grok 5
Kalshi: SpaceXAI releases Grok 4.7 before 18 September?Last trade 65¢ (bid 65¢, ask 71¢)1,785 contracts traded, 1,211 openCloses 03:59 UTC on 18 September
Kalshi: before 25 September?Bid 69¢, ask 78¢No trades yetOpened 14 September
Manifold: Grok 4.7 released by the end of September?95%13 tradersPlay money, not cash
Polymarket: no next Grok model by 30 September?17% midpointBid 1.2¢, ask 32.9¢Too thin to read

How the odds moved with Musk’s posts

Polymarket’s by-18-September contract swung between 42% and 88.5% in eleven days, and its biggest moves line up with news from Musk.

Polymarket: next Grok model released by 18 September, Yes price (UTC)
4 September, 12:00 73%
6 September, 00:00 42%
8 September, 00:00 88.5%
11 September, 12:00 81%
12 September, 00:00 65.5%
13 September, 12:00 58%
14 September, 12:00 72.5%
15 September, 12:00 64%

The sharpest fall came between 12:00 UTC on 11 September and midnight, from 81% to 65.5%. Musk posted that Grok 4.7 “needs a few more days to cook” at 17:22 UTC in that window. The price then drifted to 58% by midday on 13 September and climbed to 72.5% by midday on 14 September, after Musk’s overnight post about Grok 4.8 and his 11:19 UTC post comparing Grok 4.7 with Opus 5.0. These are twelve-hourly snapshots, so they show timing, not cause.

Every expiry so far has resolved No

Kalshi’s ladder tells the longer story. Its Grok 4.7 contracts for a release before 14 August, 21 August, 28 August, 31 August, 4 September, 8 September and 11 September have all closed at No. Octagon, which tracks the series, reported on 24 August that the before-4-September contract had already fallen from 90% to 31%. Polymarket’s by-12-September and by-14-September contracts also resolved No.

So the crowd has been too early before, and it is still pricing roughly two chances in three for this week. That is a sensible reading of Musk’s “few more days”, and also a reminder that a one-in-three chance of another miss is not small.

"On Par With Opus 5.0, Not 5.1": Reading the Capability Claim

Musk’s comparison is the most quoted part of his 14 September post, and it needs care. Anthropic’s public model overview, checked on 15 September, lists four current models by API ID: claude-fable-5-1, claude-opus-5, claude-sonnet-5 and claude-haiku-4-5. There is no public Opus 5.1. Musk may mean Anthropic’s Fable 5.1, or an Opus version he believes is coming; the post does not say.

Where “Opus 5.0 class” sits on a shared scale

The one recent table that puts these models on a single scale is in OpenAI’s GPT-6 Astra launch post, which reports the Artificial Analysis Intelligence Index v4.1.1.

Artificial Analysis Intelligence Index v4.1.1, as reported by OpenAI
Fable 5.1 65.7
Opus 5 63.1
Fable 5 62.1
GPT-6 Astra 61.2
GPT-5.6 Sol 60.9
Gemini 3.8 Flash 58.7

On that scale, “roughly on par with Opus 5.0” would put Grok 4.7 near 63: above GPT-6 Astra and Gemini 3.8 Flash, below Fable 5.1. SpaceXAI’s own Grok 4.6 launch table gave Grok 4.6 a score of 61 on the same index, but with its own test settings, so the two tables should not be merged. Taken at face value, Musk is describing a modest step up from 4.6, not the model that would “exceed all current models”, as he put it on 12 August.

Multimodal is the admitted weak spot

Musk’s own caveat is that xAI needs “to fix multimodal performance”. That is exactly the ground Gemini 3.8 Live is aimed at: speech, images and video in real time. SpaceXAI’s API model list still names no Grok 4.7, and on Musk’s description a first release would not challenge Google and OpenAI in live voice straight away.

The Third Forecast: Opus 5.2 Routing Reports

The last line of TestingCatalog’s post is the thinnest: “This week? Testers say some Claude Code Opus 5 traffic is routing to Opus 5.2 — faster and less lazy, still unofficial.” Two outlets ran with it. 36Kr published “Opus 5.2 launched late at night, is RSI really here?” on 14 September, and Biggo Finance described a “stealth rollout” the next day.

Neither cites Anthropic. At 12:00 UTC on 15 September, Anthropic’s model overview still listed claude-opus-5 as its Opus model, with no 5.1 or 5.2 identifier. Silent server-side swaps are hard to prove from outside: users notice faster, more thorough answers, but load, system prompt and caching changes produce the same impression. We treat this line as a rumour until Anthropic publishes a model ID.

Scoring the Model Forecast: Gemini 3.8 Live, Grok 4.7 and Opus 5.2

TestingCatalog’s forecast is useful because it puts a time on each claim, which means each line can be checked. This is where the three stood at midday UTC on 15 September, and what would settle them.

Forecast lineTiming givenEvidenceStatusWhat would confirm it
Gemini 3.8 Live and Live Extended Thinking“Today?”Two console screenshotsNot releasedA release note, pricing row or model page
Grok 4.7 near Opus 5.0“This week?”Musk’s posts; markets at 64% to 65% by 18 SeptemberNot releasedA SpaceXAI news post and a Grok 4.7 model in the API list
Opus 5.2 serving Claude Code traffic“This week?”Tester impressionsUnconfirmedAn Anthropic announcement or model ID

How to weigh a leak like this

The three lines carry very different weight. A quota entry is a hard artefact: someone configured it. Musk’s posts are statements of intent from the person with the most information and the weakest record on dates. Tester impressions of a faster model are the weakest evidence of the three.

Read as a set, Gemini 3.8 Live is the item most likely to exist in something close to its reported form, and also the one whose capabilities are least documented. Grok 4.7’s capabilities have been described in some detail by its owner, but its date keeps moving.

What Gemini 3.8 Live and Grok 4.7 Mean for Teams Building Voice Agents

Neither model can be used today, so the practical advice is about keeping options open rather than switching.

If you are building on Gemini Live today

Keep building on the model you can call. gemini-3.1-flash-live-preview is the current conversational option in the Gemini API, and gemini-live-2.5-flash-native-audio is the generally available one on Google Cloud. Keep the model name in configuration rather than code, because Gemini 3.8 Live will almost certainly arrive under a new identifier.

Plan a test pass for thinking behaviour too. If Gemini 3.8 Live has no minimal thinking level, the delay before the first spoken word may change, and the Extended Thinking variant may change it more. Measure time to first audio on your own call scripts before and after switching.

If you are comparing Google with OpenAI

Price real calls, not list rates. GPT-Live-1 bills session minutes, silence included, while Google bills audio in and audio out separately. Record a week of representative conversations, measure how much of each call is caller speech, model speech and silence, and apply both price lists. Then test tool calling and interruptions on your own scripts, because those are where voice agents fail in production.

If the decision is part of a wider platform choice, our AI strategy team can help structure that evaluation, and our AI models and tools hub tracks each release as it lands.

If you are waiting for Grok 4.7

Do not plan a launch around it. Musk has given four dated timelines for the model, and all of them have passed. The markets price about a two-in-three chance of a release by 18 September, which also means about a one-in-three chance of another miss. If multimodal performance is still the weak spot, voice and vision workloads will be the last to move.

The pages to watch for Gemini 3.8 Live

Four public pages should move first when Google launches Gemini 3.8 Live: the Gemini API release notes, the Gemini API pricing page, the Live API model table on Google Cloud, and the Gemini 3.8 Flash model page, whose “Gemini Live API: Not supported” line would change if the Live variant is folded into it. For Grok 4.7, the equivalents are SpaceXAI’s news page and its API model list.

Frequently Asked Questions About Gemini 3.8 Live

What is Gemini 3.8 Live?

Gemini 3.8 Live is the name of an unreleased Google model that appeared on the Google Cloud console quotas page on 15 September 2026, according to TestingCatalog. It is expected to be a real-time voice and video model for the Live API, based on Gemini 3.8 Flash. Google has not announced it.

What is Gemini 3.8 Live Extended Thinking?

It is a second model name that appeared alongside the first. Google has not explained it. The likeliest reading is a variant that reasons for longer before it answers, trading some speed for accuracy on harder requests.

When will Gemini 3.8 Live be released?

There is no date. TestingCatalog asked “Today?” on 15 September, but no launch had appeared on Google’s release notes by midday UTC. For Gemini 3.8 Flash, six days separated the first testing report from general availability.

Will Grok 4.7 come out this week?

Possibly. Polymarket priced a 64% chance of a release by 18 September and Kalshi’s before-18-September contract last traded at 65 cents at midday on 15 September. Every earlier date has been missed.

Is Grok 4.7 better than Claude Opus 5?

Musk says it will be “roughly on par with Opus 5.0”, better in some ways and worse in others, with multimodal performance still to fix. There is no independent benchmark until the model ships.

References