Tavus Griffin: the first model to pass the video Turing test — 48% of live callers thought it was human

Tavus released Griffin, billed as the world first Human Interaction Model (HIM) and the first model to pass the real-time video Turing test: in a blind study, 26 of 54 participants (48%) who had a one-minute video call with a PAL (Personified Application Layer) running on Griffin believed their partner was a real human, versus 2.4% (1 of 41) for Tavus previous stack (Phoenix 4.5 + Sparrow-2 + Raven-1). The core is full-duplex video-to-video: instead of walkie-talkie turn-taking, perception, conversational decision-making, and generation run continuously in parallel — the conversational model reassesses the exchange every sub-second mini-turn and emits control signals (what to say, how to deliver it, with what expression/stance/timing) while streaming speech and video generators follow as they arrive. The model keeps seeing and listening even while it speaks. Two engines: (1) Continuous Conversational Modeling ingests both audio and video, reading gaze and facial expression plus the user environment, so a pause for thought is not treated as end-of-turn; it can interrupt, be interrupted, and backchannel with mm-hm. (2) Audio-Visual Generation: Streaming Speech uses an autoregressive diffusion transformer (VDiT) with Tavec, a codebook-free continuous codec (48 kHz in/out, 40 values per frame at 100 frames/s, fully causal decoder with 10 ms packets, voice cloning from a 10 s clip); Streaming Video distills a large bidirectional many-step diffusion teacher in three stages (few-step DMD student, then teacher forcing into an autoregressive model, then Self-Forcing for drift-free long rollouts) into a few-step autoregressive generator with 8x VAE temporal compression — one latent equals 8 frames equals 320 ms of 720p video, averaging 0.43 s latency on H100s (half the next-fastest method). Evaluations: first on both tracks of NVIDIA VideoFDB (generation 3.83 vs human reference 3.92 and next-best 2.80, only 0.09 from human; perception 3.73 vs human 4.20 and strongest baseline MiniCPM-o 4.5 at 3.44), with best-in-class takeover-rate alignment of 62.8% / 73.8%; against four published streaming diffusion models it ranks first on DOVER, FID and THEval, second on lip-sync LSE-C at 7.27. Safety: precisely because it is hard to distinguish from a person, Griffin-Lite is limited to select trusted testers as a research preview while disclosure and safety mechanisms are developed.




