Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking are the two model names that began appearing on Google Cloud’s console quotas page on the morning of 15 September 2026, according to TestingCatalog, which posted screenshots on Threads and credited the find to researcher Bedros Pamboukian. Neither model is public. Google has not announced either name, and its developer release notes still end with the Gemini 3.8 Flash and Lyria 3.5 launches of 2 and 3 September.
The same morning, TestingCatalog folded the find into a three-line “model forecast”: Gemini 3.8 Live flagged “Today?”, Grok 4.7 flagged “This week?” and expected to land “near Opus 5.0, not 5.1”, and a tester report that some Claude Code traffic is reaching an unofficial Opus 5.2. For teams building voice products and AI agents, the first line matters most, because the newest conversational model Google offers developers is still a March preview.
This article sets out what the quota-page sighting does and does not prove, where Gemini 3.8 Live would sit in Google’s Live API lineup, how it compares with OpenAI’s GPT-Live-1, and what Elon Musk’s posts and three prediction markets say about Grok 4.7. Every price, date and probability below was read on 15 September 2026, and anything that is inference rather than published fact is labelled as such.
Table of contents
- What TestingCatalog Found: Gemini 3.8 Live on the Cloud Quota Page
- What a Quota Page Entry Proves About Gemini 3.8 Live
- Where Gemini 3.8 Live Would Fit in Google’s Live API Lineup
- Why Gemini 3.8 Live Needs Its Own Model, and What “Extended Thinking” Could Mean
- Gemini 3.8 Live Versus OpenAI’s GPT-Live-1
- How Fast Gemini 3.8 Live Could Follow the Flash Release Cycle
- The Grok 4.7 Forecast: What Musk Has Promised
- What Prediction Markets Price for Grok 4.7
- “On Par With Opus 5.0, Not 5.1”: Reading the Capability Claim
- The Third Forecast: Opus 5.2 Routing Reports
- Scoring the Model Forecast: Gemini 3.8 Live, Grok 4.7 and Opus 5.2
- What Gemini 3.8 Live and Grok 4.7 Mean for Teams Building Voice Agents
- Frequently Asked Questions About Gemini 3.8 Live
- References
What TestingCatalog Found: Gemini 3.8 Live on the Cloud Quota Page
The report is short, so it is worth reading exactly. TestingCatalog wrote that “Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking model names have started to appear on the GCP Console quotas page”, and added that “a new Gemini Live model, based on the latest Gemini 3.8 Flash, is expected to be a big leap”. The post carried two screenshots of the console list and no link to a Google source.
The two posts behind the headline
The first post is the sighting itself. A second post in the same thread supplies the context: “So far, Gemini 3.1 Live Preview is the latest Gemini Live model available via APIs.” It notes that Google recently added connector calls and Deep Research to Gemini Live, and that “OpenAI also released the GPT Live 1 model on the APIs last week, and it seems like Google has a response to that”. The credit line reads “Discovered by Bedros Pamboukian”.
A little later, TestingCatalog’s daily “MODEL FORECAST” post compressed the finding into one line: “Today? Gemini 3.8 Live and Live Extended Thinking slugs appeared on the Cloud quota page; not public.” That is where the word “slugs” in most headlines comes from. The question mark matters. It is a guess about timing, not a report that Gemini 3.8 Live has launched.
What the sighting establishes, claim by claim
Each part of the Gemini 3.8 Live report carries a different level of evidence. Here is how the claims stood at 12:00 UTC on 15 September.
| Claim | Source | Status |
|---|---|---|
| “Gemini 3.8 Live” and “Gemini 3.8 Live Extended Thinking” appear on a GCP Console quotas page | TestingCatalog screenshots | Reported with screenshots; not independently viewable |
| The models are not public | TestingCatalog forecast post | Consistent with Google’s docs: no model page, price or release note |
| Gemini 3.8 Live is based on Gemini 3.8 Flash | TestingCatalog | Inference from the name; Google has not said |
| “Expected to be a big leap” | TestingCatalog | Opinion; no benchmark published |
| A response to OpenAI’s GPT-Live-1 | TestingCatalog | Framing; the timing fits, the motive is unconfirmed |
| Launch “Today?” | TestingCatalog forecast post | No launch on any Google page by midday UTC |
What we could not verify
We could not open a Google Cloud project that shows the entries, so we cannot confirm the exact model identifiers, which quota metrics they sit under, or whether every project sees them. TestingCatalog’s wording moves between “model names” and “slugs”, and the screenshots show display names.
The identifier Google eventually ships for Gemini 3.8 Live may look different. Today’s Live models are called gemini-3.1-flash-live-preview in the Gemini API and gemini-live-2.5-flash-native-audio on Google Cloud, two quite different naming patterns, and neither puts “Live” straight after the version number the way the console entries do.
What a Quota Page Entry Proves About Gemini 3.8 Live
A quota is a limit on how much of a resource a Google Cloud project may use. Google’s quota documentation says quotas “generally apply at the Google Cloud project level”, and its model tables track limits per base model through a base_model dimension. The published table lists requests for gemini-embedding, for example, under the metric global_embed_content_input_tokens_per_minute_per_base_model.
Why model names surface on quota pages early
Before any customer can call a model at scale, Google has to create the metrics that meter it. Those entries are configuration, and configuration tends to ship before the launch blog post, the model card, the pricing row and the documentation page. That makes the quotas page a useful early signal, and it is why watchers such as TestingCatalog check it.
It is also why an entry proves less than it seems. It shows that serving infrastructure for Gemini 3.8 Live has been set up. It does not show that access has been granted, that pricing is final or that a launch date has been set. Quota entries for a model can sit unused for weeks, and a model can be renamed between configuration and launch.
A ladder of release signals
Release evidence comes in steps, from internal configuration to a public announcement. Gemini 3.8 Live has reached only the second rung.
| Signal | What it shows | Seen for Gemini 3.8 Live? |
|---|---|---|
| Internal testing reports | A model exists and is being evaluated | Implied by the name, not reported |
| Quota page entry | Serving limits have been configured | Yes, per TestingCatalog |
| Model returned by the API model list | Some keys can call it | Not reported |
| Pricing page row | Billing is defined | No, checked 15 September |
| Model documentation page | Google supports it publicly | No |
| Release notes entry | Launch | No; newest entries are 2 and 3 September |
| Launch blog post | Announcement with claims and benchmarks | No |
How quickly the last leak turned into a launch
For Gemini 3.8 Flash, the gap between the first public report of testing and general availability was short. Business Insider reported on 27 August that Google employees were already testing the model, and Google made gemini-3.8-flash generally available on 2 September, six days later.
That is one data point, not a rule, and a quota entry is a different kind of signal from staff testing. A Live model also has more to prove than a text model before launch: latency, turn-taking, interruptions and audio quality all have to hold up in real conversations, not just in a benchmark harness. Gemini 3.8 Live could follow in days or sit in configuration for weeks.
Where Gemini 3.8 Live Would Fit in Google's Live API Lineup
The Live API is Google’s interface for real-time conversation. Google describes it as a stateful connection over WebSockets that takes in a continuous stream of audio, video frames or text and replies with audio or text. The models behind it have their own names and release cadence, separate from the Flash and Pro text models.
Under the hood, these models fold speech recognition, natural language processing and speech generation into one model instead of a pipeline of three, which is where most of the latency saving comes from. Gemini 3.8 Live would be the first Live model to carry the 3.8 version number.
The Live models Google lists today
Google’s two developer platforms list different Live models, which is part of why a new name is notable. This is the lineup Gemini 3.8 Live would join.
| Model | Platform | Status | What it does |
|---|---|---|---|
| gemini-3.1-flash-live-preview | Gemini API | Preview, released 26 March 2026 | Audio-to-audio dialogue; text, image, audio and video in; text and audio out |
| gemini-live-2.5-flash-native-audio | Google Cloud | Generally available, marked “Recommended” | Low-latency voice agents with affective dialog, proactive audio and tool use |
| gemini-2.5-flash-native-audio-preview-12-2025 | Gemini API | Preview, released 12 December 2025 | Earlier native audio model for the Live API |
| gemini-3.5-transcribe-live-preview | Google Cloud | Preview | Speech-to-text only; returns text and does not hold a conversation |
| gemini-3.5-live-translate-preview | Gemini API | Preview | Real-time speech-to-speech translation in 70+ languages |
| Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking | Console quotas page only | Not public | Unknown |
What Gemini 3.1 Flash Live can and cannot do
The current preview sets the bar that Gemini 3.8 Live has to clear. Its model page lists an input limit of 131,072 tokens and an output limit of 65,536, with function calling, Google Search grounding and thinking supported. Structured outputs, code execution and caching are not supported.
The page also lists gaps. Function calling is synchronous only, so the model “will not start responding until you’ve sent the tool response”. And “proactive audio and affective dialogue” are “not yet supported in Gemini 3.1 Flash Live”. The older 2.5 native audio model on Google Cloud does support both, so developers currently choose between a newer reasoning core and richer conversational behaviour. A Gemini 3.8 Live that closed that gap would matter more than any single benchmark.
What Google charges for Live today
Google has published no price for Gemini 3.8 Live. These are the paid-tier list prices for the Live models it does sell, per million tokens in US dollars, with Google’s own per-minute equivalents in brackets.
| Model | Audio input | Audio output | Text input | Text output |
|---|---|---|---|---|
| Gemini 3.1 Flash Live Preview | $3.00 ($0.005/min) | $12.00 ($0.018/min) | $0.75 | $4.50 |
| Gemini 2.5 Flash Native Audio | $3.00 | $12.00 | $0.50 | $2.00 |
| Gemini 3.5 Live Translate | $3.50 ($0.0053/min) | $21.00 ($0.0315/min) | Not listed | Not listed |
| Gemini 3.8 Live | Not published | Not published | Not published | Not published |
Google notes that output prices include thinking tokens. If an Extended Thinking variant of Gemini 3.8 Live spends more of each turn reasoning, the audio rate could stay the same while the cost of each turn rises.
Why Gemini 3.8 Live Needs Its Own Model, and What "Extended Thinking" Could Mean
There is a plain reason Google cannot simply point the Live API at the model it launched on 2 September. The Gemini 3.8 Flash page on Google Cloud lists “Gemini Live API: Not supported”. The text model takes text, image, audio and video in, with a 1,048,576-token context window and up to 65,536 output tokens, but it answers in text. A live voice model needs audio output, streaming turn detection and interruption handling, which is separate training and serving work. That is the gap Gemini 3.8 Live would fill.
Gemini 3.8 Flash has no minimal thinking level
The more telling clue is how thinking works. Gemini 3.1 Flash Live uses thinking levels of minimal, low, medium and high, and Google says “the default is minimal to optimize for lowest latency”.
Gemini 3.8 Flash is different. Google’s documentation says thinking_level MINIMAL “is not available for 3.8 Flash” and that setting it “will return an API validation error”. The supported values are low, medium and high. If Gemini 3.8 Live inherits that floor, even its fastest setting would think more than today’s Live default, and a voice model that thinks longer answers later.
Reading the “Extended Thinking” name
This part is inference, because Google has said nothing. The plainest reading is that Gemini 3.8 Live is tuned for conversational speed, and Gemini 3.8 Live Extended Thinking allows longer reasoning per turn for tasks such as working through a support case, checking a calculation or planning a multi-step action while the caller waits. Voice products already live with that trade-off: a slow answer feels broken, but a fast wrong answer is worse.
The label is also unusual for Google. Its API talks about “thinking levels” and, before that, “thinking budgets”, while “extended thinking” is the phrase Anthropic uses for Claude. Quota-page names are internal, so the Gemini 3.8 Live Extended Thinking name may not survive to launch in that form.
Two variants, two jobs
If the two names do map to a fast and a deliberate model, developers would route calls between them in much the same way they already choose thinking levels. This is how the split could look, stated as a working hypothesis rather than a specification.
| Need | Likely fit | Why |
|---|---|---|
| Receptionist, booking, simple FAQs | Gemini 3.8 Live | Turn latency matters more than depth |
| Technical support with diagnosis | Gemini 3.8 Live Extended Thinking | Wrong answers cost more than a pause |
| Tutoring and worked problems | Extended Thinking | Checking steps is the product |
| Live translation | Neither; Gemini 3.5 Live Translate exists | A dedicated model is already on sale |
Gemini 3.8 Live Versus OpenAI's GPT-Live-1
TestingCatalog framed the sighting as Google’s answer to OpenAI, and the timing supports that reading. OpenAI brought GPT-Live-1 to its API on 10 September 2026, 64 days after GPT-Live launched in ChatGPT on 8 July, and we covered that launch in detail in our GPT-Live-1 API article. Five days later, the Gemini 3.8 Live names appeared in Google’s console.
What is known on each side
The fair comparison today is three-way: OpenAI’s shipped model, Google’s shipped preview, and the unreleased Gemini 3.8 Live.
| Factor | GPT-Live-1 | Gemini 3.1 Flash Live Preview | Gemini 3.8 Live |
|---|---|---|---|
| Status | Generally available in the API since 10 September 2026 | Preview since 26 March 2026 | Names on a quotas page |
| Interface | /v1/live/sessions | Live API over WebSockets | Presumably the Live API |
| List price | $0.05 per minute of session time, billed per second | $0.005/min audio in, $0.018/min audio out | Not published |
| Context | 128,000 tokens | 131,072 input tokens | Unknown |
| Harder reasoning | Can hand work to a separate reasoning model; four of OpenAI’s seven launch charts used one | Built-in thinking levels, default minimal | A separate Extended Thinking name exists |
| Published benchmarks | Yes, including tool calling and interruptions | Model card | None |
The per-minute arithmetic
The two companies price differently, so a like-for-like figure needs an assumption. Take a ten-minute call in which the app streams the caller’s microphone for all ten minutes and the model speaks for four of them.
At list rates, GPT-Live-1 costs 10 × $0.05 = $0.50, because OpenAI bills session time, silence included. Gemini 3.1 Flash Live costs 10 × $0.005 + 4 × $0.018 = $0.05 + $0.072 = $0.122, before thinking tokens, tool calls or growing context. That is about a quarter of OpenAI’s list price on this assumption, but the gap will move once Google prices Gemini 3.8 Live.
For reference, Gemini 3.8 Flash launched at $0.75 per million input tokens and $3.75 per million output tokens, the same introductory rate as 3.7 Flash, and both figures double to $1.50 and $7.50 from 1 January 2027. A newer core could well carry a higher Live rate than the March preview.
Why the list price will not decide it
A cheaper rate does not settle the choice. OpenAI published measured results for GPT-Live-1 on tool calling and interruption handling, and customer results on real call volumes. Google has published nothing for Gemini 3.8 Live, and its current preview still lacks asynchronous function calling. Until Google publishes, the comparison that matters in production is GPT-Live-1 against the models developers can actually call.
How Fast Gemini 3.8 Live Could Follow the Flash Release Cycle
Google’s Flash line has moved quickly this summer. Gemini 3.6 Flash arrived on 21 July, 3.7 Flash on 13 August and 3.8 Flash on 2 September, gaps of 23 and 20 days; we covered the latest in our Gemini 3.8 Flash in Agent Studio article. The conversational Live line has not kept pace. The newest Live model for dialogue in the Gemini API, gemini-3.1-flash-live-preview, dates from 26 March.
Counted to 15 September, the gap between the text models and the Live model is stark.
What the cadence suggests, and what it does not
A Live model built on 3.8 Flash would close a gap of almost six months between Google’s conversational and text models. Nothing in the cadence dates the Gemini 3.8 Live launch, though. Google has not tied Live releases to Flash releases: the Live model numbered 3.1 has served while three Flash generations shipped, and the 2.5 native audio model went through dated previews in September and December 2025 before a version reached general availability on Google Cloud.
Where “Today?” stood at midday
TestingCatalog’s forecast asked whether Gemini 3.8 Live would ship on 15 September. At 12:00 UTC that day, the Gemini API release notes still showed Lyria 3.5 on 3 September as the newest entry, Google Cloud’s release notes had nothing newer on Live, and neither pricing page mentioned the model. A same-day launch was still possible, but none of the public surfaces had moved.
The Grok 4.7 Forecast: What Musk Has Promised
The second line of the forecast turns to SpaceXAI. TestingCatalog wrote: “This week? Grok 4.7 should land near Opus 5.0, not 5.1; multimodal still needs work.” That wording tracks a post Elon Musk made on 14 September. Unlike Gemini 3.8 Live, where the evidence is a console entry, the evidence for Grok 4.7 is almost entirely what one person has posted, and every date he has given has passed.
The posts that set expectations
These are the five Musk posts that set the Grok 4.7 timeline, with their view counts as read through the fxtwitter API at midday on 15 September.
| Date (UTC) | What Musk wrote | Views |
|---|---|---|
| 24 July | “Grok 4.6 in 2 weeks and Grok 4.7 in 4 weeks” | 1.80 million |
| 12 August | “significantly better than 4.6 and should be ready in 3 to 4 weeks” | 1.33 million |
| 2 September | “Grok 4.7 comes out in 10 days” | 10.29 million |
| 11 September | “needs a few more days to cook” | 2.05 million |
| 14 September | “roughly on par with Opus 5.0, not 5.1. Better in some ways, worse in others. We need to fix multimodal performance.” | 1.99 million |
Why the model slipped, in Musk’s words
The 11 September post is the only one that gives a technical reason. Musk wrote that xAI “might have penalized response length too much (or something) in RL, as it still gives up on hard tasks (that it can do!) too early and isn’t yet sufficiently rigorous in checking its work”.
In plain terms, reinforcement learning taught the model to keep its answers short, and it learned to stop before finishing difficult work. That kind of tuning problem can be fixed in days or weeks, which is why the forecast still says “this week”. We set out the full sequence of dates, and the 2.5-trillion-parameter Grok 4.8 Musk says comes next, in our Grok 4.8 article. Counted to 15 September, Grok 4.7 is 25 days past the four-week estimate of 24 July.
What Prediction Markets Price for Grok 4.7
Three venues let people bet on the release date. At midday on 15 September they broadly agreed that Grok 4.7 is more likely than not to ship within days, while disagreeing on how sure to be. None of them offers a contract on Gemini 3.8 Live.
Three venues, three different questions
The contracts are not identical, so the prices should not be averaged. The table below sets out what each one actually asks.
| Venue and question | Price at 12:00 UTC, 15 September | Activity | Catch |
|---|---|---|---|
| Polymarket: next Grok model (4.7 or higher) released by 18 September? | 64% (bid 63¢, ask 65¢) | $15,172 traded | Any 4.7-or-higher Grok counts, including Grok 5 |
| Kalshi: SpaceXAI releases Grok 4.7 before 18 September? | Last trade 65¢ (bid 65¢, ask 71¢) | 1,785 contracts traded, 1,211 open | Closes 03:59 UTC on 18 September |
| Kalshi: before 25 September? | Bid 69¢, ask 78¢ | No trades yet | Opened 14 September |
| Manifold: Grok 4.7 released by the end of September? | 95% | 13 traders | Play money, not cash |
| Polymarket: no next Grok model by 30 September? | 17% midpoint | Bid 1.2¢, ask 32.9¢ | Too thin to read |
How the odds moved with Musk’s posts
Polymarket’s by-18-September contract swung between 42% and 88.5% in eleven days, and its biggest moves line up with news from Musk.
The sharpest fall came between 12:00 UTC on 11 September and midnight, from 81% to 65.5%. Musk posted that Grok 4.7 “needs a few more days to cook” at 17:22 UTC in that window. The price then drifted to 58% by midday on 13 September and climbed to 72.5% by midday on 14 September, after Musk’s overnight post about Grok 4.8 and his 11:19 UTC post comparing Grok 4.7 with Opus 5.0. These are twelve-hourly snapshots, so they show timing, not cause.
Every expiry so far has resolved No
Kalshi’s ladder tells the longer story. Its Grok 4.7 contracts for a release before 14 August, 21 August, 28 August, 31 August, 4 September, 8 September and 11 September have all closed at No. Octagon, which tracks the series, reported on 24 August that the before-4-September contract had already fallen from 90% to 31%. Polymarket’s by-12-September and by-14-September contracts also resolved No.
So the crowd has been too early before, and it is still pricing roughly two chances in three for this week. That is a sensible reading of Musk’s “few more days”, and also a reminder that a one-in-three chance of another miss is not small.
"On Par With Opus 5.0, Not 5.1": Reading the Capability Claim
Musk’s comparison is the most quoted part of his 14 September post, and it needs care. Anthropic’s public model overview, checked on 15 September, lists four current models by API ID: claude-fable-5-1, claude-opus-5, claude-sonnet-5 and claude-haiku-4-5. There is no public Opus 5.1. Musk may mean Anthropic’s Fable 5.1, or an Opus version he believes is coming; the post does not say.
Where “Opus 5.0 class” sits on a shared scale
The one recent table that puts these models on a single scale is in OpenAI’s GPT-6 Astra launch post, which reports the Artificial Analysis Intelligence Index v4.1.1.
On that scale, “roughly on par with Opus 5.0” would put Grok 4.7 near 63: above GPT-6 Astra and Gemini 3.8 Flash, below Fable 5.1. SpaceXAI’s own Grok 4.6 launch table gave Grok 4.6 a score of 61 on the same index, but with its own test settings, so the two tables should not be merged. Taken at face value, Musk is describing a modest step up from 4.6, not the model that would “exceed all current models”, as he put it on 12 August.
Multimodal is the admitted weak spot
Musk’s own caveat is that xAI needs “to fix multimodal performance”. That is exactly the ground Gemini 3.8 Live is aimed at: speech, images and video in real time. SpaceXAI’s API model list still names no Grok 4.7, and on Musk’s description a first release would not challenge Google and OpenAI in live voice straight away.
The Third Forecast: Opus 5.2 Routing Reports
The last line of TestingCatalog’s post is the thinnest: “This week? Testers say some Claude Code Opus 5 traffic is routing to Opus 5.2 — faster and less lazy, still unofficial.” Two outlets ran with it. 36Kr published “Opus 5.2 launched late at night, is RSI really here?” on 14 September, and Biggo Finance described a “stealth rollout” the next day.
Neither cites Anthropic. At 12:00 UTC on 15 September, Anthropic’s model overview still listed claude-opus-5 as its Opus model, with no 5.1 or 5.2 identifier. Silent server-side swaps are hard to prove from outside: users notice faster, more thorough answers, but load, system prompt and caching changes produce the same impression. We treat this line as a rumour until Anthropic publishes a model ID.
Scoring the Model Forecast: Gemini 3.8 Live, Grok 4.7 and Opus 5.2
TestingCatalog’s forecast is useful because it puts a time on each claim, which means each line can be checked. This is where the three stood at midday UTC on 15 September, and what would settle them.
| Forecast line | Timing given | Evidence | Status | What would confirm it |
|---|---|---|---|---|
| Gemini 3.8 Live and Live Extended Thinking | “Today?” | Two console screenshots | Not released | A release note, pricing row or model page |
| Grok 4.7 near Opus 5.0 | “This week?” | Musk’s posts; markets at 64% to 65% by 18 September | Not released | A SpaceXAI news post and a Grok 4.7 model in the API list |
| Opus 5.2 serving Claude Code traffic | “This week?” | Tester impressions | Unconfirmed | An Anthropic announcement or model ID |
How to weigh a leak like this
The three lines carry very different weight. A quota entry is a hard artefact: someone configured it. Musk’s posts are statements of intent from the person with the most information and the weakest record on dates. Tester impressions of a faster model are the weakest evidence of the three.
Read as a set, Gemini 3.8 Live is the item most likely to exist in something close to its reported form, and also the one whose capabilities are least documented. Grok 4.7’s capabilities have been described in some detail by its owner, but its date keeps moving.
What Gemini 3.8 Live and Grok 4.7 Mean for Teams Building Voice Agents
Neither model can be used today, so the practical advice is about keeping options open rather than switching.
If you are building on Gemini Live today
Keep building on the model you can call. gemini-3.1-flash-live-preview is the current conversational option in the Gemini API, and gemini-live-2.5-flash-native-audio is the generally available one on Google Cloud. Keep the model name in configuration rather than code, because Gemini 3.8 Live will almost certainly arrive under a new identifier.
Plan a test pass for thinking behaviour too. If Gemini 3.8 Live has no minimal thinking level, the delay before the first spoken word may change, and the Extended Thinking variant may change it more. Measure time to first audio on your own call scripts before and after switching.
If you are comparing Google with OpenAI
Price real calls, not list rates. GPT-Live-1 bills session minutes, silence included, while Google bills audio in and audio out separately. Record a week of representative conversations, measure how much of each call is caller speech, model speech and silence, and apply both price lists. Then test tool calling and interruptions on your own scripts, because those are where voice agents fail in production.
If the decision is part of a wider platform choice, our AI strategy team can help structure that evaluation, and our AI models and tools hub tracks each release as it lands.
If you are waiting for Grok 4.7
Do not plan a launch around it. Musk has given four dated timelines for the model, and all of them have passed. The markets price about a two-in-three chance of a release by 18 September, which also means about a one-in-three chance of another miss. If multimodal performance is still the weak spot, voice and vision workloads will be the last to move.
The pages to watch for Gemini 3.8 Live
Four public pages should move first when Google launches Gemini 3.8 Live: the Gemini API release notes, the Gemini API pricing page, the Live API model table on Google Cloud, and the Gemini 3.8 Flash model page, whose “Gemini Live API: Not supported” line would change if the Live variant is folded into it. For Grok 4.7, the equivalents are SpaceXAI’s news page and its API model list.
Frequently Asked Questions About Gemini 3.8 Live
What is Gemini 3.8 Live?
Gemini 3.8 Live is the name of an unreleased Google model that appeared on the Google Cloud console quotas page on 15 September 2026, according to TestingCatalog. It is expected to be a real-time voice and video model for the Live API, based on Gemini 3.8 Flash. Google has not announced it.
What is Gemini 3.8 Live Extended Thinking?
It is a second model name that appeared alongside the first. Google has not explained it. The likeliest reading is a variant that reasons for longer before it answers, trading some speed for accuracy on harder requests.
When will Gemini 3.8 Live be released?
There is no date. TestingCatalog asked “Today?” on 15 September, but no launch had appeared on Google’s release notes by midday UTC. For Gemini 3.8 Flash, six days separated the first testing report from general availability.
Will Grok 4.7 come out this week?
Possibly. Polymarket priced a 64% chance of a release by 18 September and Kalshi’s before-18-September contract last traded at 65 cents at midday on 15 September. Every earlier date has been missed.
Is Grok 4.7 better than Claude Opus 5?
Musk says it will be “roughly on par with Opus 5.0”, better in some ways and worse in others, with multimodal performance still to fix. There is no independent benchmark until the model ships.
References
TestingCatalog: Gemini 3.8 Live model names on the GCP Console quotas page
TestingCatalog: Model Forecast, 15 September 2026
Gemini 3.1 Flash Live Preview (Gemini API documentation)
Gemini Live API overview (Google Cloud documentation)
Gemini 3.8 Flash model page (Google Cloud documentation)
Generative AI on Agent Platform quotas and system limits
Gemini 3.8 Flash and 3.8 Flash Cyber (Google)
Gemini Live adds Deep Research as Notebooks come to AI Mode (9to5Google)
Build more natural voice experiences with GPT-Live-1 in the API (OpenAI)
GPT-Live 1 model page (OpenAI API documentation)
Elon Musk on X: Grok 4.7 roughly on par with Opus 5.0
Elon Musk on X: Grok 4.7 needs a few more days to cook
Polymarket: Next Grok Model (4.7+) released by…?
Manifold: Will Grok 4.7 be released by the end of September 2026?
Musk Timeline Pushes Grok 4.7 Early Release Odds (Octagon)
Introducing Grok 4.6 (SpaceXAI)
More AI coverage: explore Progressive Robot's AI Models, Tools & Releases hub — hands-on reviews, setup guides and benchmarks in one place.