Tavus says its Griffin model fooled 26 of 54 people on one-minute video calls
NVIDIA's own leaderboard backs the benchmark half of the claim, and Tavus is holding the model back over deception.
Tavus, a San Francisco company that builds conversational video AI, announced Griffin on October 1, 2026, and says it's the first model to pass a real-time video Turing test. In the company's own study, 26 of 54 people who had a one-minute video call with it came away thinking they'd talked to a real person.
That's 48%. Tavus's previous system, run through the same test, fooled 1 of 41. You can't use Griffin yet: a research preview called Griffin-Lite is going to a small group of testers, and customers aren't getting it.
- Maker
- Tavus
- Model
- Griffin-Lite, a research preview
- Live study
- 26 of 54 participants believed their partner was a real person
- Previous system
- 1 of 41
- Benchmark
- NVIDIA VideoFDB
- Who can use it
- select trusted testers
How the test worked
People were recruited through an independent research platform and told they'd be matched with another participant for a one-minute call about what they were looking forward to this year. The other participant was Griffin, generating her face, voice and replies as the call went. Only at the end of the survey were they asked whether it had crossed their mind that their partner might not be a real person. Then everyone was told.
Over half said the thought never came up during the call. The ones who did suspect usually did it early, within the first 20 seconds, according to Tavus.
Until now, almost no one has said yes.
It's a small sample, and it's the company's own (Tavus ran the study, wrote it up and picked the format). Participants weren't told an AI might be on the call, so nobody went in trying to catch it out. I'd want to see the same test run on people who'd been warned.
What NVIDIA's leaderboard shows
The benchmark claim is easier to check, because the scores sit on NVIDIA's own page. VideoFDB lists Tavus Griffin Lite first on both of its tracks. On perception, which asks whether a model reads the moment from a person's audio and video, it scores 3.73 out of 5. Gemini 2.5 Flash Native scores 3.17, OpenAI's gpt-realtime 2.75, and the human reference 4.20.
The generation track is a thinner field. Griffin's 3.83 sits close to the human 3.92, but only two other systems are on that table, and both are Gemini 2.5 driving a separate avatar. On picking the right nonverbal cue it still trails people, 2.83 to 3.18. Its median response time there is 1,892 milliseconds; for the humans it's 900.
What Tavus says about the risk
Tavus's post names the obvious problem itself.
The same properties that make Human Interaction Models powerful interfaces for natural communications between human and machine allow them to deceive a human into believing it is not AI.
The company says it's working on disclosure features and with AI safety groups, and that it expects to release Griffin "very soon after these safety concerns are addressed." Tavus didn't give a date. For now, testers can request access through a form on the Griffin page.