Artificial Intelligence

MAI Realtime full duplex speech illustrated as two identical smooth solids on a dark platform joined by two separate channels of light running in opposite directions at the same moment, both lit along their whole length, a conversation in which both sides speak and listen simultaneously

MAI Realtime: Microsoft Tests Its First Bidirectional Voice Model

On 2 August 2026 a model called MAI Realtime surfaced as a hidden early-access entry in Microsoft’s MAI Playground, described in reporting as a bidirectional full-duplex system with two voices, seventeen languages, two selectable turn-taking modes and partner access already granted. Microsoft has confirmed none of it: no model card, no benchmark, no price, no availability date. This guide separates what was actually reported from what is inference, explains why full duplex is an architectural change rather than a latency improvement, and sets out where MAI Realtime would sit against MAI-Voice-2, MAI-Transcribe-1.5 and the GPT-Realtime model currently doing the speech-to-speech work inside Azure’s Voice Live API. It covers the OpenAI dependency this closes, what the existing MAI price points imply, the 200-millisecond turn-taking budget that decides whether a voice agent feels human, where voice deployments actually fail in production, the EU AI Act transparency obligations that became applicable the same day, voice as biometric data under GDPR, the TCPA position on synthetic voices, and a ninety-day plan whose value does not depend on MAI Realtime ever shipping.

Read more
Concentric sound rings stretching into a blank panel and settling as evenly spaced bars, speech becoming editable text in the Meta AI prompt box

Meta Is Building Voice Dictation Into the Meta AI Website: What Changes When You Can Talk to the Prompt Box

Voice dictation is the one item on Meta’s leaked Meta AI web feature list that still has not shipped, while Research, Reasoning, Presentation and scheduled tasks all landed. That gap is the story: Meta has a full speech stack, a 1,600-language ASR system and dictation on Quest, yet the meta.ai prompt box still only accepts typing. Here is what Meta is actually building, why voice dictation is a different feature from voice mode, how browser dictation gets implemented, and the privacy question a microphone on a web page raises.

Read more
CHAT