Griffin Lite is the first public look at what Tavus calls a Human Interaction Model: an AI that watches, listens and talks back on a live video call, with a generated face, voice and body, all at once. Tavus announced it on 1 October 2026 as a research preview for a select group of testers, and says a more powerful model, Griffin, will follow once safety concerns are addressed.

The headline claim is striking. In a Tavus study, 26 of 54 people who had a one-minute video call with a Griffin Lite character believed they had spoken to a real person. Tavus’s previous system fooled 1 of 41. The company says that makes it the first model to pass a real-time video Turing test, and it is unusually frank that the same ability could be used to deceive.

This article explains what Griffin Lite is, how it works, what the benchmarks do and do not show, what changed on the Tavus developer platform in September, and what businesses should think about before putting a lifelike AI face in front of customers. For background on voice cloning risks, see our report on a Japanese anime actor’s fight with TikTok over AI voice cloning.

What Tavus Announced With Griffin Lite

griffin lite tavus real time video ai preview b life size cardboard standee of a person

The announcement is a long research post on the Tavus website, signed by co-founder and chief executive Hassaan Raza and head of research Ioannis Patras, dated 1 October 2026 in San Francisco.

A new class of model

Tavus describes Griffin as “the world’s first Human Interaction Model (HIM), a new class of model designed to understand and generate face-to-face real-time human interaction.” In plain terms, it is one system that sees the person on camera, hears them, decides when to speak, and generates the reply as video and audio together, rather than a chain of separate tools.

Who can use Griffin Lite today

Almost nobody yet. Tavus says Griffin Lite “will not be available for use for customers at this time, though it is available for select trusted testers as a research preview.” There is a request form. The page adds: “Griffin isn’t on the Tavus platform yet. It’ll come once we’ve worked out how to release it safely.”

Who helped build it

Tavus thanks Baseten, Daily and Cerebrium for the infrastructure behind the preview, Queen Mary University of London for research collaboration, and NVIDIA for building and scoring the benchmark it leans on most heavily.

How Griffin Lite Works: Video In, Video Out

griffin lite tavus real time video ai preview c lucky cat waving one raised paw

Most real-time AI avatars today work as what Tavus calls a cascade: one system transcribes your speech, a language model writes a reply, a voice engine speaks it, and a face model animates it. Each hand-off adds delay and loses information, such as your tone or what is on camera. Griffin Lite is built to avoid that chain.

Continuous conversation, not turns

Griffin Lite makes a decision at regular sub-second intervals, which Tavus calls mini-turns. At each one it decides whether to stay quiet, nod, say “mm-hm”, take the turn, or stop because you interrupted. It reads both audio and video, so a pause while you look away is not treated as the end of your sentence.

Streaming speech

The voice side is a fast autoregressive diffusion transformer that generates speech piece by piece as instructions arrive, rather than waiting for a full sentence. It uses a codec Tavus calls Tavec, which turns 48 kHz audio into 40 numbers per 10-millisecond frame, 100 frames a second. Audio leaves in packets as small as 10 ms, and Tavus says the system can clone a voice from about 10 seconds of audio.

Streaming video

The video side generates 720p video in 320-millisecond chunks in real time from a single reference photo. At 25 frames a second, each chunk is eight frames. Tavus says it controls the whole scene, including arms, fingers, the chair, shadows and the background, not just the face. It distilled a large, slow diffusion model into a fast, few-step one in three stages to make that possible.

Part of Griffin LiteWhat Tavus says it does
Conversational modelReads audio and video, decides at sub-second intervals whether to speak, react or wait
Speech generatorStreams speech in packets as small as 10 ms; clones a voice from about 10 seconds
Tavec codec48 kHz audio; 40 values per frame, 100 frames per second; 6,000 frames per minute
Video generator720p in 320 ms chunks from one reference image; three diffusion steps per chunk
PerceptionUses what it sees, such as a Rubik’s cube or a shared screen

What it looks like in practice

Tavus’s demonstration clips show Griffin Lite coaching someone through a Rubik’s cube while watching it turn, playing Simon Says and calling out a bluff, telling a story with a tester while they talk over each other, and guiding someone through soldering a circuit board, speaking up when the next step is due rather than whenever the person goes quiet.

The Benchmarks Behind Griffin Lite

griffin lite tavus real time video ai preview d badminton rackets and a shuttlecock over a net

Tavus ran three kinds of test: the video generator alone, the full system on an NVIDIA benchmark, and live calls with people. Each tells a different part of the story.

The video Turing test

Participants were recruited through an independent research platform and told they would have a one-minute call with another participant about what they were looking forward to this year. Their partner was a Griffin Lite character. Afterwards they were asked whether their partner was a real person. 26 of 54 said yes, which is 48%. On Tavus’s previous stack (Phoenix-4.5, Sparrow-2 and Raven-1), 1 of 41 said yes, which is 2.4%.

Share of participants who thought the AI was a real person after a one-minute call (Tavus study)

Griffin Lite, 26 of 54: 48%
Previous Tavus system, 1 of 41: 2.4%

Those who said “real” averaged 79% confidence, and those who said “AI” averaged 81%. More than half said the possibility never crossed their mind during the call. Those who did suspect usually did so within the first 20 seconds. On a seven-point scale, participants rated Griffin Lite 5.4 for seeming natural, 5.6 for trustworthiness and 5.8 for wanting to talk to it again. Conversation flow scored lowest, at 4.9.

NVIDIA’s VideoFDB benchmark

VideoFDB is NVIDIA’s benchmark for full-duplex audio-visual conversation, and Tavus says NVIDIA ran the scoring independently in September 2026 with its own judge. On the generation track, which scores the model’s own speech and video, Griffin Lite scored 3.83 out of 5 against a human reference of 3.92. The next-best system, Gemini 2.5 with Anam, scored 2.80.

On the perception track, which tests whether the model understands what it sees and hears, Griffin Lite scored 3.73 against a human reference of 4.20, ahead of MiniCPM-o 4.5 at 3.44, Gemini 2.5 Flash Native at 3.17 and OpenAI’s gpt-realtime at 2.97. Tavus says that put it first among 15 models.

VideoFDB perception track, judge score out of 5 (bar length = score ÷ 5)

Human reference: 4.20
Griffin Lite: 3.73
MiniCPM-o 4.5 (audio-only): 3.44
Gemini 2.5 Flash Native: 3.17
OpenAI gpt-realtime (audio-only): 2.97

Where the “37% ahead” and “12x closer” figures come from

Tavus’s page says Griffin Lite is “37% ahead of the next best AI at reacting in the moment.” That matches the generation scores: 3.83 minus 2.80 is 1.03, and 1.03 divided by 2.80 is 36.8%. The claim of being “more than 12x closer to human performance” also checks out: Griffin Lite is 0.09 below the human score, while the next system is 1.12 below (3.92 minus 2.80), and 1.12 divided by 0.09 is about 12.4.

Speed and picture quality

Tested against four published streaming video models, Griffin Lite’s video generator averaged 0.43 seconds from audio arriving to its effect appearing on screen, on NVIDIA H100 chips, which Tavus says is half the next fastest. It ranked first on three video quality measures (DOVER, FID and THEval) and second on lip sync (LSE-C), at 7.27.

Reading the results with care

These are strong numbers, but they come with caveats. The Turing study is small, the calls lasted one minute, and Tavus ran it. Participants were told they were meeting another participant, which primes them to expect a human. VideoFDB scores come from a language-model judge. None of this makes the results wrong, but it means Griffin Lite has not yet been tested by outsiders in long, real conversations.

Tavus Updates Its Developer Platform

griffin lite tavus real time video ai preview e sloth hanging from a branch

The other half of the news is the platform developers can use today. Tavus says 150,000 developers and businesses already build what it calls PALs, AI characters you talk to face to face, on its Phoenix, Raven and Sparrow models. Its documentation changelog shows a busy September.

MCP connectors

On 18 September Tavus added MCP connectors, which let a PAL use apps such as Gmail, Google Calendar, Slack and Linear, or any third-party MCP server, to send an email, book a meeting, post to a channel or file a ticket during a conversation.

Memory Stores

On 16 September it launched Memory Stores, giving each PAL a persistent memory of every participant. Tavus builds an evolving profile and timeline automatically, and developers can pin specific facts.

Microsoft Teams, Google Meet and Zoom

On 12 September PALs gained the ability to join Microsoft Teams calls, in addition to Google Meet and Zoom. A PAL can be invited to a calendar event with its own email address.

Faces, voices and languages

Phoenix-4.5, the newest face model, arrived on 2 September and became the default for new faces on 9 September. It handles more natural movement below the neck, glasses and jewellery, and animated styles. Tavus added 100 stock voices on 21 August and 16 more on 9 September, each speaking all 42 supported languages, plus a tavus-auto engine that picks the best speech provider for each voice and language.

Date (2026)Tavus platform change
21 August100 stock voices, custom voices, tavus-auto TTS, Sparrow-2 opt-in
28 AugustSSO, team accounts, simpler PAL publishing
2 SeptemberPhoenix-4.5 face model
12 SeptemberMicrosoft Teams support
16 SeptemberMemory Stores
18 SeptemberDynamic greetings and MCP connectors
24 SeptemberAccount-scoped S3 recording keys
1 OctoberGriffin Lite research preview (testers only)

Why the platform work matters

The September changes are the plumbing a lifelike agent needs: memory so it remembers you, connectors so it can act, meeting integrations so it can turn up where work happens. When Griffin Lite or its successor reaches the platform, it will inherit all of it. That is why the developer updates and the model preview belong in the same story.

Questions Griffin Lite Still Has to Answer

griffin lite tavus real time video ai preview f clay head on a sculptors modelling stand

The preview is impressive on paper, but several practical questions will decide whether Griffin Lite, or the full Griffin model, becomes something businesses actually use.

How long can it hold a conversation?

The Turing study used one-minute calls. Real customer conversations, tutoring sessions and interviews run for 10 to 60 minutes. Tavus says it trained the video generator specifically so that long rollouts do not drift, but it has not published results for long sessions. Until it does, nobody outside the tester group knows how Griffin Lite behaves in minute twenty.

What will it cost?

Tavus has not published pricing. Generating 720p video and speech in real time on H100-class chips is expensive, and Tavus already bills its existing platform by conversation minutes. Expect Griffin Lite, or its successor, to cost more per minute than the current Phoenix stack, at least at first.

How will disclosure work?

Tavus says it is building “safe disclosure features” but has not said what they are. A label on screen, a spoken statement at the start of a call, or a watermark in the video would each work differently, and regulators in the EU expect disclosure that people actually notice.

Can outsiders check the results?

NVIDIA’s VideoFDB scores are published on its leaderboard, which gives the benchmark claims some independent footing. The Turing study is Tavus’s own. Independent testers and researchers will want to repeat it with longer calls and with participants who know an AI might be on the other end. Those results will say more about Griffin Lite than any launch post.

The Safety Question: A Face That Passes for Human

Tavus does not hide the risk. Its post says: “The same properties that make Human Interaction Models powerful interfaces for natural communications between human and machine allow them to deceive a human into believing it is not AI.”

What Tavus says it is doing

The company says further alignment and safety work is needed before a safe release, that it is building “safe disclosure features”, and that it is working with organisations focused on AI safety. Restricting Griffin Lite to trusted testers is part of that approach. The full Griffin model will only arrive “very soon after these safety concerns are addressed.”

Deepfake fraud is already real

Video-call impersonation is not hypothetical. In 2024, staff at the engineering firm Arup in Hong Kong were tricked into transferring about $25 million after a video call with deepfaked colleagues. Voice cloning scams keep growing too; see our report on an AI messaging scam that cost Intesa millions. A model that half of people mistake for a human in a minute raises the stakes.

The disclosure rules

In the EU, the AI Act’s transparency duties in Article 50 require that people are told when they are talking to an AI system, unless it is obvious, and that deepfakes are labelled. Those duties apply from 2 August 2026. The UK has no AI-specific law yet, but consumer protection, data protection and fraud law all apply. Our article on California’s synthetic performer disclosure law covers the US direction of travel.

Why the Turing result cuts both ways

A Turing pass sounds like a product feature, and for tutoring or customer service it partly is. But the study worked because people were not told. In any real deployment, honest disclosure is both the law in many places and good practice, and once people know they are talking to an AI, the “passes for human” number matters less than whether the conversation is genuinely useful.

Where Griffin Lite Fits in Real-Time Video AI

Griffin Lite enters a crowded field of real-time avatars and voice agents, and the VideoFDB leaderboard names several of the rivals.

Cascaded avatars

Companies such as Anam, which appears on the leaderboard paired with Gemini 2.5, build talking faces on top of separate speech and language models. HeyGen and D-ID offer interactive avatars for sales and support. These systems are available today and are improving fast, but they are the cascade design that Tavus argues Griffin Lite moves beyond.

Voice-first models

OpenAI’s gpt-realtime and Google’s native Gemini audio models are scored on VideoFDB in audio-only form. They are excellent conversationalists, but they do not generate a face. Google has moved toward video with a live avatar for Gemini 3.8 Live; see our report on Gemini 3.8 Live.

ApproachExamplesGenerates a face?Available now?
Unified video-to-videoGriffin LiteYes, full sceneTesters only
Cascaded avatarTavus Phoenix stack, Anam, HeyGen, D-IDYesYes
Speech-to-speechOpenAI gpt-realtime, Gemini native audioNoYes
Open modelsMiniCPM-o 4.5No (scored audio-only)Yes

What Griffin Lite Means for Businesses

You cannot buy Griffin Lite today, but it shows where real-time video AI is heading within the next year, so it is worth planning for.

Where it could help

Tavus suggests tutors that notice confusion, rehearsal partners for difficult conversations, and support agents that can look at a broken part held up to the camera. Those are reasonable uses, especially for training. Interview practice and onboarding are obvious early fits.

Where to be careful

Sales and support are where deception risk is highest. If customers cannot tell whether they are speaking to a person, complaints and regulatory problems follow. Any deployment should disclose clearly, record consent for any likeness or voice used, and give people an easy route to a human.

Practical steps now

Start with the cascaded tools that exist, measure whether video adds anything over voice for your use case, and write a disclosure and consent policy before you need it. Our AI strategy team helps businesses decide where agents add value, and our IT security specialists help staff spot deepfake video calls. If you plan to build on the Tavus API, the computer vision and speech pieces are the parts to test hardest with real customers.

Griffin Lite FAQ

What is Griffin Lite?

Griffin Lite is a research preview of Tavus’s first Human Interaction Model, a single AI system that sees, hears and responds on a live video call by generating a person’s face, voice and movement in real time.

Can I use Griffin Lite?

Not yet, unless you are among the select testers. Tavus says it is not available to customers and is not on its platform. There is a request form for testers.

Did Griffin Lite pass the Turing test?

In Tavus’s own study, 26 of 54 people (48%) believed a Griffin Lite character was human after a one-minute video call. That is a strong result, but it comes from a small, short, company-run study.

How fast is it?

Tavus says its video generator averages 0.43 seconds from audio to on-screen effect on H100 chips, generates 720p video in 320-millisecond chunks, and streams speech in packets as small as 10 milliseconds.

What changed on the Tavus developer platform?

In September Tavus added MCP connectors, Memory Stores, Microsoft Teams support, the Phoenix-4.5 face model, more voices across 42 languages, dynamic greetings and account-scoped recording storage.

Is it safe?

Tavus itself says the model could deceive people into thinking it is human and is holding back a wider release until disclosure and safety features are ready.

References and Further Reading