Power Your Software Testing with AI Agents and Cloud
The Native AI-Agentic Cloud Platform to Supercharge Quality Engineering. Test Intelligently and Ship Faster.
- TestMu AI (Formerly LambdaTest)
- /
- Blog
- /
- Tavus Griffin: What Its Benchmarks Mean for Video Agent Testing
Tavus Griffin: What Its Benchmarks Mean for Video Agent Testing
Tavus Griffin fooled 48% of people in a one-minute video call. See what its NVIDIA VideoFDB scores show about where video agents fail and what to test first.
Published on:
Fifty-four people joined a one-minute video call expecting to chat with another study participant about what they were looking forward to this year. Their partner was an AI persona running on Tavus Griffin, generating her face, voice, and replies live, and 26 of them, 48%, said afterward they had talked with a real person, according to Tavus's Griffin announcement. On Tavus's previous system, 2.4% did.
Tavus calls Griffin the first Human Interaction Model, and NVIDIA's independent VideoFDB benchmark puts its preview version at the top of both of its leaderboards. Outside a small group of trusted testers, nobody can use it yet.
If you ship an on-camera agent, such as an AI interviewer or a video support agent, those results matter before you can use Griffin at all. The benchmark breakdown shows where the top-scoring system still trails people, and each of those gaps maps to a check your own video agent tests should include.
TL;DR
Tavus Griffin is a Human Interaction Model, a full-duplex video-to-video AI model that Tavus announced on 1 October 2026. It perceives a person's audio and video and generates its own voice, face, and body in real time. Its preview version, Griffin-Lite, is limited to select trusted testers.
- Availability: Can you use Tavus Griffin today? No. Griffin-Lite is a research preview for trusted testers while Tavus builds disclosure features. Tavus's current Phoenix-4.5, Raven-1, and Sparrow-2 models remain available.
- Turing test: Did Griffin pass a video Turing test? Tavus says yes: 26 of 54 participants (48%) took Griffin-Lite for a real person after a one-minute call, against 2.4% for its previous system.
- Benchmarks: On NVIDIA's VideoFDB, Griffin-Lite scored 3.83 out of 5 on generation, 0.09 below the human reference, and 3.73 on perception, the top non-human score on both tracks.
- Response time: Does Griffin respond as fast as a human? No. Griffin-Lite's median VideoFDB latency was 1,892 ms on generation against 900 ms for the human reference.
- Testing implications: Griffin's weakest areas, response time, nonverbal cues, and conversational flow, are the first things to test in any video agent, and TestMu AI Agent Testing grades them from a recorded session.
What Is Tavus Griffin?
Tavus Griffin is a full-duplex, video-to-video model that Tavus announced on 1 October 2026, in a post by co-founder and CEO Hassaan Raza and head of research Ioannis Patras. Tavus defines a Human Interaction Model as one designed to understand and generate face-to-face, real-time human interaction: it listens while it talks and attends to expressions and pauses as well as words.
Most real-time video agents work as a cascade: speech recognition transcribes you, a language model replies, and separate systems turn the reply into a voice and a face. Every handoff adds delay and drops signals, such as your tone or what is on camera, that the next system never sees. Griffin folds perception, turn decisions, and speech and video generation into one system that runs them all at once.
Tavus's announcement shows these capabilities in recorded conversations with people who had never used Griffin:
- Behavior and emotion - it laughs, changes its tone, and shifts its expression and gestures in response to what the other person says and does.
- Full-scene generation - it generates every pixel of every frame from one reference image, including hands, the chair, shadows, and the background.
- Full-duplex conversation - it can interrupt, back-channel with an "mm-hm", or stop the moment you cut in, without losing its place.
- Perception - it uses what it sees, for example coaching someone through a Rubik's Cube while watching the cube turn.
- Temporal understanding - it tracks how long a silence has lasted and speaks up when the next step of a task is due.
Griffin's Two-Engine Architecture
Griffin splits the work between two engines that run concurrently for the whole conversation:
- Continuous Conversational Modeling - perceives the incoming audio and video, decides when and how to respond, and emits control signals for what to say plus emotional tone, stance, facial expression, and gesture.
- Audio-Visual Generation - a streaming speech generator and a streaming video generator that turn those control signals into voice and video as they arrive.
The conversational engine makes its decisions at regular sub-second intervals instead of once per turn. That lets it read a pause for thought as thinking and keep waiting, and nod or back-channel while you are still speaking. Because it perceives video, it can also use your gaze, your expression, and anything you show the camera.
The generation side is built for streaming:
- Speech - an autoregressive diffusion transformer generates speech one latent chunk at a time, so it starts talking before the sentence is complete, and it can clone a voice from about 10 seconds of audio.
- Audio codec - Tavus's Tavec codec maps 48 kHz audio to 40 values per frame at 100 frames per second and streams audio packets as small as 10 ms.
- Video - a few-step autoregressive diffusion generator, distilled in three stages from a large many-step model, produces 720p video in 320 ms chunks in real time.
- Steering - streaming controls let the conversational engine push a gesture, a glance away, or a change of emotion into the very next chunk of video.
For testing, the architecture decides where failures show up. A cascade adds delay at every handoff and cannot react while the user is talking, while a full-duplex model can react at any moment, including the wrong one, so test turn-taking in both directions: when the agent should speak up and when it should stay quiet.
Tavus Griffin on NVIDIA's VideoFDB Benchmark
VideoFDB is NVIDIA's benchmark for full-duplex audio-visual conversation, built from 237 annotated clips of real two-person video calls that span 11 nonverbal conversational dynamics, such as turn-taking, backchannels, and gaze aversion. A language-model judge scores each response from 0 to 5 on two tracks: perception, for whether the agent reads the situation, and generation, for whether its own audio-visual output fits.
These are the overall scores on NVIDIA's leaderboard, where Griffin-Lite is the only entry listed on both tracks. The highest non-human score on each track is in bold:
| System | Generation (0 to 5) | Perception (0 to 5) |
|---|---|---|
| Human reference recordings | 3.92 | 4.20 |
| Tavus Griffin-Lite | 3.83 | 3.73 |
| Gemini 2.5 + Anam (cascaded avatar) | 2.80 | Not on this track |
| Gemini 2.5 + Keyframe (cascaded avatar) | 2.39 | Not on this track |
| MiniCPM-o 4.5 (audio-only) | Not on this track | 3.44 |
| MiniCPM-o 4.5 (audio and video) | Not on this track | 3.40 |
| Gemini 2.5 Flash Native | Not on this track | 3.17 |
| Gemini 3.1 Flash Live (audio-only) | Not on this track | 3.03 |
| OpenAI gpt-realtime (audio-only) | Not on this track | 2.97 |
On generation, Griffin-Lite finished 1.03 points ahead of the next system and 0.09 points below the human reference. On perception it led the strongest baseline by 0.29 points, the top score among the 15 entries evaluated.
NVIDIA also publishes the rubric scores behind each track, and they show where the remaining gap to humans sits. In the timing rows, the first figure is takeover-rate (TOR) alignment, how closely a system's choices about when to speak match the reference conversations, and the second is median latency:
| Track and rubric | Griffin-Lite | Human reference | Best other entry |
|---|---|---|---|
| Generation: fluency | 4.25 | 4.42 | 3.48 (Gemini 2.5 + Anam) |
| Generation: dyadic affect | 4.40 | 4.14 | 3.21 (Gemini 2.5 + Anam) |
| Generation: nonverbal cue appropriateness | 2.83 | 3.18 | 1.71 (Gemini 2.5 + Anam) |
| Generation: timing | 62.8% / 1,892 ms | 78% / 900 ms | 44% / 2,840 ms (Gemini 2.5 + Anam) |
| Perception: fluency | 3.60 | 4.16 | 3.45 (MiniCPM-o 4.5, audio-only) |
| Perception: conversational flow | 3.67 | 4.20 | 3.76 (MiniCPM-o 4.5, audio-only) |
| Perception: visual grounding | 3.92 | 4.24 | 3.63 (MiniCPM-o 4.5, audio and video) |
| Perception: timing | 73.8% / 2,232 ms | 90% / 1,400 ms | 73% / 720 ms (MiniCPM-o 4.5, audio and video) |
The breakdown separates what Griffin already does well from what still lags:
- Emotional matching - dyadic affect is the one rubric where Griffin-Lite scores above the human reference, 4.40 against 4.14.
- Nonverbal cues - nonverbal cue appropriateness is its lowest generation score, 2.83 against 3.18 for humans, though still well ahead of the cascaded avatars.
- Response time - its median latency runs about a second behind the human reference on generation and 0.8 seconds behind on perception.
- Conversational flow - on perception, MiniCPM-o 4.5 in its audio-only configuration scores higher on flow, 3.76 against 3.67.
NVIDIA's own findings explain the size of Griffin's lead over cascades. Audio-driven avatars move only in response to the speech they produce, so they cannot add a nod or an expression while the user is talking, and the two cascaded systems in the benchmark had median latencies of 2.8 to 3.5 seconds.
Both tracks are scored by a model rather than by human raters, the same LLM as a judge approach many teams use to grade their own agents, so read small score differences with care.
How Fast Is Griffin's Video Generator?
Tavus also tested Griffin-Lite's video generator on its own, against four published streaming diffusion models that each turn speech and a reference image into a talking face. Tavus reports these results:
- Latency - true audio-to-video latency averaged 0.43 seconds on H100s, half that of the next fastest method, because the generator produces one latent at a time with no lookahead audio.
- Visual quality - first among the five models on DOVER and FID, the two standard video quality measures, and on THEval, a recently published framework built for talking heads.
- Lip sync - second on LSE-C, at 7.27. Tavus notes that LSE-C rewards pronounced mouth movement, including movement past the point where it looks natural, which THEval accounts for.
These comparisons come from Tavus rather than an independent evaluator, so weigh them below the VideoFDB results. In a conversation, the lower latency matters most when a person interrupts or a nod is due, because the face reacts sooner.
Did Tavus Griffin Pass a Video Turing Test?
Tavus recruited participants through an independent research platform and told them they would be matched with another participant for a one-minute call about what they were looking forward to this year. Only at the end of a survey were they asked whether it had crossed their mind that their partner might not be a real person. Every participant was then told the partner had been an AI.
Tavus ran the same protocol on its previous stack of Phoenix-4.5, Sparrow-2, and Raven-1:
| Measure | Griffin-Lite | Phoenix-4.5 + Sparrow-2 + Raven-1 |
|---|---|---|
| Participants | 54 | 41 |
| Believed the partner was a real person | 26 (48%) | 1 (2.4%) |
| Average confidence, "real person" answers | 79% | Not reported |
| Average confidence, "AI" answers | 81% | Not reported |
Participants also rated Griffin-Lite on a 7-point scale:
- Seemed natural - 5.4.
- Seemed trustworthy - 5.6.
- Would enjoy talking again - 5.8, holding at 5.4 among participants who said it was an AI.
- Was really listening - 5.5.
- Conversation flowed naturally - 4.9, the lowest of the five.
Over half of the participants said the possibility never crossed their mind during the call, and those who did suspect tended to do so within the first 20 seconds.
Tavus designed and ran the study, the sample was small, and each call lasted one minute on a single light topic. Treat the headline result as a strong early signal, and note that flow, the rubric where an audio-only model beat Griffin-Lite on VideoFDB, also rated lowest with people.
Note: In Tavus's study, conversation flow rated lowest. TestMu AI Agent Testing scores conversation flow, question handling, response quality, and avatar presentation from 0 to 10 on a recorded video session, so you can see which one your agent is losing. Try TestMu AI free!
When Can You Use Tavus Griffin?
Not yet. Griffin-Lite is a research preview for select trusted testers, Tavus says it will not be available to customers for now, and a wider release of a more powerful model is to follow.
Tavus ties the release to safety work:
- Deception risk - Tavus says the properties that make Human Interaction Models natural also let them deceive a person into believing they are not talking to an AI.
- Disclosure features - it is building safe disclosure features and working with organizations tackling AI safety before release.
- Trusted-tester access - teams can request preview access through a form linked from the announcement.
- Available today - developers can build on Phoenix-4.5 for rendering, Raven-1 for perception, and Sparrow-2 for turn-taking, models Tavus says 150,000 developers and businesses already use.
The disclosure question reaches beyond Tavus. If your video agent could pass for a person, test that it identifies itself as an AI when asked directly and wherever your policies require it.
How to Turn Griffin's Benchmarks Into Video Agent Tests
A transcript cannot show whether an agent talked over a user, froze mid-sentence, or drifted out of lip-sync, and those are the areas Griffin's own scores point to. If the top-scoring system still falls short there, your agent probably does too, so I would map each finding to a test:
| Benchmark finding | What to test in your video agent |
|---|---|
| Griffin-Lite's median latency runs about a second behind the human reference on VideoFDB | Response timing: measure the gap between the end of a user's turn and the agent's reply on your own infrastructure and network. |
| Nonverbal cue appropriateness is Griffin-Lite's lowest generation rubric | Nonverbal behavior: check that expressions and gestures fit the moment, including while the user is still talking. |
| An audio-only model beat Griffin-Lite on perception flow, and flow rated lowest in Tavus's study | Turn-taking under pressure: interrupt the agent, pause mid-thought, and go quiet, then check that it yields, waits, and resumes without losing its place. |
| LSE-C rewards exaggerated mouth movement, so a strong lip-sync score can still look unnatural | Lip-sync and presentation: review lip-sync, facial motion, and audio and video quality on the recording itself. |
| Participants rated Griffin-Lite 5.6 out of 7 for seeming trustworthy | Task accuracy: a convincing face makes a wrong answer more believable, so grade the substance of every answer against the task. |
| Tavus is holding Griffin back to build disclosure features | AI disclosure: confirm the agent says it is an AI when a user asks. |
| The human study used one-minute calls on a single topic, and VideoFDB uses a model as judge | Varied users: run each scenario with blunt, chatty, distracted, and impatient users, and repeat it, because a live conversation never plays out the same way twice. |
Agent Testing runs those checks on a live session. TestMu AI joins your agent's joinable web session URL as a simulated candidate with a real face and voice, holds an improvised conversation, and records it, with no SDK, no phone number, and no change to the agent.
- Personas and fan-out - one scenario expands across avatar faces, personas such as blunt, chatty, distracted, or impatient, test profiles, and up to 10 iterations, with up to 50 sessions per suite run.
- Success criteria - you write the pass conditions, the agent passes only if every criterion is met, and anything that cannot be verified from the recording counts as not met.
- Evidence - each run returns a downloadable recording, a turn-labeled transcript, a criteria table with Expected, Achieved, Evidence, and Confidence columns, and timestamped evidence.
- Fair scoring - the simulated candidate counts as test equipment, so its own glitches cannot lower your agent's result, and sessions that fail on the TestMu AI side are marked Inconclusive.
Video testing is enabled per organization, sessions run two at a time and last 30 to 600 seconds, and agents that exist only inside Zoom, Meet, or Teams are not supported yet. Point runs at a staging URL, because every session is a real conversation with whatever sits behind the link.
For the full method, see video agent testing and video simulation testing, and for a practitioner view, the Testμ 2026 session on testing video AI agents.
Getting Started With Tavus Griffin
If you need Griffin before its wider release, request trusted-tester access from Tavus, and keep building on Tavus's current models or your existing video stack until then. Write the success criteria and personas for your agent now, run them against today's version, and keep the results to rerun when you switch models; the guide to getting started with Agent Testing covers the setup.
Author
Anubhav Singhmaar is an AI Product Manager at TestMu AI driving Kane CLI, the command-line tool that brings browser automation to the terminal, turning natural-language flows into runs in a real Chrome browser that return pass or fail with shareable proof. He owns the roadmap and prioritization and works with engineering to ship developer-facing features. Before TestMu AI, he spent over four years at Sprinklr owning enterprise voice AI across APAC and EMEA. A mechanical engineer turned product manager, he grounds guidance in real QA workflows.
Reviewer
Sirajuddin Khan is Vice President of Product Management at TestMu AI (formerly LambdaTest), where he drives the company's agentic AI product strategy, building a suite of autonomous agents that includes Agentic Browsers and Agentic Visual Testing and shifting the unit of work from test execution to autonomous outcomes. One of the company's earliest product leaders, he has owned the roadmap for the high-performance execution cloud and grew the cross-browser testing products from early adoption to market leadership. He brings over a decade of experience across SaaS, B2B, and eCommerce, with earlier product roles at Wydr and ShopClues, where his catalog and search work cut delivery SLAs and lifted seller activity. Sirajuddin holds an MBA in Information Technology from Sikkim Manipal University and a B.Tech in Computer Science Engineering from Maharshi Dayanand University.
Tavus Griffin FAQs
Did you find this page helpful?
More Related Blogs
TestMu AI forEnterprise
Get access to solutions built on Enterprise
grade security, privacy, & compliance
- Advanced access controls
- Advanced data retention rules
- Advanced Local Testing
- Premium Support options
- Early access to beta features
- Private Slack Channel
- Unlimited Manual Accessibility DevTools Tests





