PAPER DEEP DIVE
DyaPlex: Full-Duplex Speech-Motion Model for Dyadic Interaction
We present DyaPlex, a streaming, full-duplex speech-and-motion model designed for dyadic interaction. To capture the continuous and reciprocal nature of human communication, this full-duplex capability empowers the agent to simultaneously perceive and generate both speech and physical motion in a streaming fashion. At its core, our method leverages the strong priors of a foundational full-duplex speech model and integrates a novel motion pathway, thereby achieving fully synchronized multi-modal interaction. Specifically, we design a dual-tower Transformer architecture that preserves the zero-shot conversational reasoning of a frozen base speech model while constructing a deeply coupled, streaming motion pathway. By introducing a unified dyadic token interleaving mechanism and guiding cross-attention via a time-aligned speech-motion RoPE, our model effectively aligns autoregressive motions with rich latent speech features. Trained on the 4,000-hour Seamless Interaction dataset, our model effectively captures cross-speaker dependencies and establishes new state-of-the-art performance across both monadic and dyadic human interaction benchmarks.
Paper metadata
Title: DyaPlex: Full-Duplex Speech-Motion Model for Dyadic Interaction
Authors: Koki Nagano and Hongyu Liu (co-first authors; Liu was an intern from HKUST), Seonwook Park, Tianye Li, Amrita Mazumdar, Christian Jacobsen, Shengze Wang, Michael Stengel, Rajarshi Roy, Ka Chun Cheung, Simon See, Shalini De Mello (NVIDIA)
Links: arXiv:2606.03874 (https://arxiv.org/abs/2606.03874 , v1, June 2 2026); project page https://research.nvidia.com/labs/amri/projects/DyaPlex/
Code and data status: no public code is released with the paper; the project page currently carries only the arXiv link and the author list. The upstream pieces needed to reproduce are open: the speech tower PersonaPlex lives at github.com/NVIDIA/personaplex, the motion tokenizer is adapted from GestureLSM (arXiv:2501.18898), and both baselines are released at github.com/facebookresearch/audio2photoreal and github.com/ziqiaopeng/DualTalk. Training data is the public Seamless Interaction dataset.
Figure 1: the causal full-duplex setting. Top row is the partner's speech and motion input; bottom row is the agent listening while speaking and producing backchanneling motion; the right panel lists the two target applications, human-agent/robot dyadic interaction and synthetic dyadic speech-motion data generation (paper Figure 1).
One-sentence summary
DyaPlex freezes the full-duplex speech model PersonaPlex as a speech tower, hangs a trainable 32-layer motion tower off it, and binds both speakers' speech and full-body motion onto one causal timeline via dyadic token interleaving plus time-aligned speech-motion RoPE, becoming the first model that can listen and speak while watching the partner's body and streaming its own, cutting monadic gesture FGD from the baseline's 57 to 5.6 (times 1e-3) on 4,000 hours of Seamless data.
Background and motivation: what dyadic interaction lacks is not fidelity, it is simultaneity
In natural dyadic interaction, listening and expressing are never discrete alternating turns but a highly synchronized streaming process in which speech and body movement are present at the same time. Full-duplex speech models in the Moshi and PersonaPlex lineage have made fluent verbal exchange possible, but genuine social presence demands tight coupling of speech and motion: a listener keeps nodding and mirroring the speaker's posture without breaking the speech flow, and the speaker reads those motion responses and adjusts in real time. This mutual speech-motion perception is the foundation next-generation Embodied Conversational Agents cannot avoid building.
A capability table (Table 1) makes the gap concrete. Audio2Photoreal (A2P) and DyaDiT rely on non-causal diffusion models: high visual quality but offline-only generation, structurally excluded from streaming interaction. ViBES uses an LLM backbone to jointly model speech and motion but pairs it with a non-causal motion tokenizer and stays completely blind to the partner's body. Among concurrent causal methods, SARAH streams dyadic motion yet perceives only the partner's 2D floor-projected position, throwing away the semantics of actual gestures; MIBURI achieves full-duplex interaction in the speech domain but its motion generation remains monadic, since it cannot receive partner motion as input and must condition the agent's actions on speech alone.
The missing piece is a unified causal perception-generation loop, and its absence breaks communication itself. The paper's example is precise: subconscious mirroring and silent backchanneling. In natural encounters a listener continuously copies the speaker's posture or supplies sustained motion feedback such as vigorous nodding without interrupting speech, while the speaker perceives those motion responses and dynamically adjusts their own behavior, closing the loop. Frameworks forced to act as delayed turn-takers, or conditioned on speech alone, are entirely oblivious to these concurrent motion cues, and they break down completely in scenarios that require perceiving the partner's speech and motion at the same time.
There is also a data-scale gap. Prior corpora in Table 1 range from 8 hours (A2P custom), 50 hours (SARAH's Embody 3D, DualTalk custom), 70 hours (MIBURI's BEAT2) and 182 hours (DyaDiT's Seamless subset) up to 1,000 hours (ViBES's Converse 3D). DyaPlex trains on roughly 4,000 hours of Seamless Interaction, 20 to 500 times larger. The high-order dependencies of two-person coordination, who yields the turn and who mirrors whom, are exactly what small corpora cannot teach.
So the design goal is stated in one line: condition the agent's actions on concurrent multi-modal cues, speech and motion, at the frame level. The implementation is a division of labor between two towers. The frozen speech tower preserves PersonaPlex's zero-shot conversational reasoning and exposes per-layer residual-stream hidden states, injected into a trainable motion tower through cross-attention. The motion tower embeds both participants' motions in a single autoregressive sequence so partner and agent motion features interact deeply under self-attention, while time-aligned RoPE injects explicit relative temporal distances into cross-attention and stops the cross-modal mapping from degenerating into time-agnostic fixed speech-feature retrieval. Both towers are strictly causal, so streaming generation is an architectural property rather than a retrofit.
Preliminaries: two frozen foundations
The PersonaPlex speech tower. DyaPlex reuses PersonaPlex, a causal Transformer built on Moshi, with hidden dimension $d_{s}=4096$ and $L_{s}=32$ layers. Both speakers' speech is tokenized at $f_{s}=12.5$ Hz with the Mimi neural speech codec; at each frame PersonaPlex receives a 17-way interleaving of text and dyadic speech codebooks, and the embedding layer sums those 17 per-codebook embeddings into a single $d_s$-dimensional vector per frame. PersonaPlex is frozen throughout training. The crucial property is that for every transformer block $\ell=1,\dots,L_{s}$ it exposes its post-block residual-stream hidden states $\mathcal{H}_{\ell}\in\mathbb{R}^{T\times d_{s}}$, hierarchical features encoding linguistic and prosodic nuance that the motion tower consumes directly.
The body-part-based RVQ-VAE motion tokenizer. The motion side adapts the part-aware RVQ-VAE of GestureLSM, retrained end-to-end on Seamless and modified into a causal streaming architecture: the encoder consumes 25 fps motion and emits tokens at the speech-aligned rate $f_{m}=12.5$ fps, while the decoder applies 2x temporal upsampling to reconstruct 25 fps SMPL-X output. Four independent decoders specialize in upper body, hands, lower body and face. Each frame is encoded by $K=22$ codes (18 body plus 4 face) drawn from a shared vocabulary of size $V_{\text{mot}}=4096$, partitioned into four disjoint 1024-entry bands. At inference the codes are routed to their respective decoders to reconstruct SMPL-X and FLAME parameters. Encoder and decoder are both frozen during motion-tower training.
Why discrete tokens are the right representation. Quantizing motion into integer codes drops motion generation straight into the language-model next-token paradigm: causal sampling, streaming output and KV reuse all come for free, while the per-part banded vocabulary plus a training-time band-mask keeps probability mass off structurally invalid cross-part tokens. The price is decoding 22 codes one at a time per frame, a cost the authors themselves acknowledge in their limitations.
Method: one interleaved sequence, two kinds of attention
Notation. Consider a dyad $(A,B)$ where, without loss of generality, $A$ is the partner and $B$ the agent. Their Mimi audio tokens are $\mathbf{s}^{A}_{1:T},\mathbf{s}^{B}_{1:T}$ at $f_{s}=12.5$ Hz and their RVQ-VAE motion tokens $\mathbf{m}^{A}_{1:T},\mathbf{m}^{B}_{1:T}$ at $f_{m}=12.5$ fps. Each motion frame is a $K=22$-dimensional integer vector $\mathbf{m}_{t}=(c^{(1)}_{t},\dots,c^{(K)}_{t})$ over a shared vocabulary of size $V_{\text{mot}}=4096$.
Speech hidden-state extraction with a hybrid input prefix. The motion tower conditions on PersonaPlex's per-layer residual-stream hidden states $\{\mathcal{H}_{\ell}\}_{\ell=1}^{L_{s}}$ via cross-attention at every block. Exposing all 32 layers rather than a single final embedding gives cross-attention access to every intermediate speech representation, and motion-tower blocks are paired one-to-one with PersonaPlex layers so cross-attention can learn its own hierarchical speech representations. Because PersonaPlex is causal, training precomputes all hidden states in a single teacher-forced forward pass; at inference the states are produced autoregressively on the fly by the standard PersonaPlex inference loop. One distribution-matching detail: PersonaPlex needs a hybrid text-and-voice system prompt to initialize autoregressive generation, so at training time a constructed system prompt is explicitly prepended to each Seamless clip before hidden states are extracted.
Motion tower architecture. The motion tower is a causal decoder-only Transformer with $L_{m}=32$ blocks (one per PersonaPlex layer), dimension $d_{m}=1024$, $h_{m}=16$ self-attention heads of head dimension $d_{m}/h_{m}=64$, RoPE self-attention and SwiGLU feed-forward layers. A single shared embedding $\mathbf{E}\in\mathbb{R}^{V\times d_{m}}$ maps token ids to features, where the vocabulary $V=V_{\text{mot}}+2$ includes two special speaker tags $\mathtt{[A]}$ and $\mathtt{[B]}$.
Self-attention with dyadic interleaving. The two motion streams are flattened into a single token sequence that alternates speaker tags and their $K$ RVQ codes at every frame:
$$\mathbf{M}=\big[\,\underbrace{\mathtt{[A]},\,\mathbf{m}^{A}_{1},\,\mathtt{[B]},\,\mathbf{m}^{B}_{1}}_{\text{frame 0}},\;\underbrace{\mathtt{[A]},\,\mathbf{m}^{A}_{2},\,\mathtt{[B]},\,\mathbf{m}^{B}_{2}}_{\text{frame 1}},\;\dots\,\big],\qquad L_{\text{step}}=2(K+1)=46.$$
The per-frame step length is therefore $L_{\text{step}}=46$ tokens. Causal self-attention over the unified sequence $\mathbf{M}$ simultaneously models three distinct dependencies: intra-frame coherence among one speaker's $K$ RVQ codes, within-frame cross-speaker reactions from $A$ to $B$, and long-range temporal dynamics across consecutive frames. This is the first architecture in which both sides of a dyadic conversation share a single autoregressive motion prior, and the P-FD and Delta-User columns of Table 2 show that adding partner motion improves dyadic motion quality by an order of magnitude.
Cross-attention with time-aligned speech-motion RoPE. Each motion-tower block $\ell$ is paired one-to-one with the corresponding PersonaPlex block and interacts through a multi-head cross-attention sub-layer with $h_{c}=12$ heads of per-head dimension $d_{c}=64$; queries come from the interleaved motion stream and keys and values from the frozen speech hidden states $\mathcal{H}_{\ell}$. Concretely, the motion-tower hidden state $\mathbf{h}_{t}\in\mathbb{R}^{d_{m}}$ at flattened position $t$ is projected by a trainable matrix $\mathbf{W}_{q}^{\ell}\in\mathbb{R}^{d_{m}\times h_{c}d_{c}}$ into the query $\mathbf{q}_{t}$, while trainable matrices $\mathbf{W}_{k}^{\ell},\mathbf{W}_{v}^{\ell}\in\mathbb{R}^{d_{s}\times h_{c}d_{c}}$ project $\mathcal{H}_{\ell}$ into keys $\mathbf{K}_{\ell}$ and values $\mathbf{V}_{\ell}$. The learned projections matter because PersonaPlex's features were natively optimized for audio synthesis; the motion tower needs the capacity to remap them into representations suited to gesture generation.
The temporal alignment is the cleanest move in the paper. Every motion token is assigned a query position equal to its actual frame index:
$$q_{\text{pos}}(t)\;=\;\big\lfloor t\,/\,L_{\text{step}}\big\rfloor,$$
so all 46 tokens comprising a single motion frame share exactly the same query position. Rotary positional embeddings are then applied to queries and keys at those positions:
$$\tilde{\mathbf{q}}_{t}=\mathrm{RoPE}\big(\mathbf{q}_{t},\,q_{\text{pos}}(t)\big),\qquad \tilde{\mathbf{k}}_{s}=\mathrm{RoPE}\big(\mathbf{K}_{\ell}[s],\,s\big),$$
where $s$ is the temporal index of the speech tokens. Since the motion sampling rate matches the speech sampling rate ($f_{m}=f_{s}$), $q_{\text{pos}}(t)$ and $s$ live on a unified temporal axis. RoPE computes attention strictly from the relative offset $q_{\text{pos}}(t)-s$, so the ideal solution reduces to a diagonal alignment in which each motion frame attends directly to its concurrent speech frame: an inductive bias written into the structure, which the network learns to exploit. Without RoPE, cross-attention has no explicit positional signal and must recover alignment implicitly from the motion query and speech key alone, yielding suboptimal speech-motion alignment; the attention maps in Appendix B (Figure 4 here) make this visible, with attention smeared without RoPE and lit up along the diagonal with it.
Speech context window and causality. To guarantee strict causality for real-time streaming, cross-attention dictates that a motion token $t$ may attend only to speech frames at or preceding its concurrent motion frame, enforced by a causal mask:
$$M_{t,s}=\begin{cases}1,&\text{if } s\leq q_{\text{pos}}(t)\\ 0,&\text{otherwise}\end{cases}$$
In principle cross-attention can access an indefinite speech history without extra motion self-attention overhead; for simplicity the paper aligns the speech context with the motion tower's 4096-token window, bounding the receptive field to 89 frames, about 7.1 seconds at 12.5 fps.
Training objective. The motion tower parameters $\theta$ are optimized by teacher-forced next-token prediction over the interleaved sequence, conditioned on the precomputed speech states $\{\mathcal{H}_{\ell}\}$. To stop the model assigning probability mass to structurally invalid tokens across the four disjoint RVQ codebooks, a band-mask sets out-of-band logits to $-\infty$ before the softmax. The masked cross-entropy is
$$\mathcal{L}_{\text{CE}}(\theta)=-\sum_{t\in\mathcal{S}}\log p_{\theta}(x_{t+1}\mid x_{\leq t};\{\mathcal{H}_{\ell}\}),\qquad \mathcal{L}=\mathcal{L}_{\text{CE}}+\beta\,\mathcal{L}_{\text{VA}},\quad \beta=0.01,$$
where $\mathcal{S}$ denotes the 18 supervised body code positions per frame, using the official SMPL-H body data (body plus hands) shipped with Seamless. The paper focuses on body motion generation and "Ours" refers to this body-only base model; a separate variant that also predicts face codes is used only for the qualitative demos in Figure 1 and the supplementary video. Following MIBURI, a linear voice-activation head predicts the binary speaking/listening state $v_{t}$ at each valid code position, giving an auxiliary binary cross-entropy loss $\mathcal{L}_{\text{VA}}$.
Streaming inference. At inference the system is a causal streaming sampler: the PersonaPlex speech tower synthesizes agent B's speech while producing the per-layer hidden states $\mathcal{H}_{\ell}$ that encode the dyadic conversational context. Conditioned on $\mathcal{H}_{\ell}$, the motion tower either generates both speakers' motion jointly (both-speaker mode, for synthetic dyadic data generation) or only the agent's motion (agent-only mode, for human-agent or robot interaction). In the latter, given an observed partner-motion prefix $\mathbf{m}^{A}_{1:f}$ and PersonaPlex hidden states for both speakers up to frame $f$, the observed partner tokens are filled into the $\mathtt{[A]}$ slots and sampling happens exclusively at the agent's $\mathtt{[B]}$ slots:
$$\hat{c}^{(k)}_{f}\,\sim\,\mathrm{topk}\!\big(\mathrm{softmax}(\mathrm{logits}(\mathbf{x}_{<t})/\tau),\;K_{\text{top}}\big),$$
where $\mathbf{x}_{<t}$ is the partial interleaved sequence up to flat position $t$. Speaker tags are inserted deterministically and the ground-truth partner motion $\mathbf{m}^{A}_{f}$ is copied into its designated positions to preserve the dyadic structure; autoregressive sampling is restricted to the $\mathtt{[B]}$ code positions with temperature $\tau=1.0$ and $K_{\text{top}}=200$. Because the speech hidden states are continuously produced by a frozen streaming speech tower and the part-aware RVQ decoders are inherently causal, the entire pipeline from partner audio input through PersonaPlex and the motion tower down to final SMPL-X reconstruction maintains strict causality and can be executed chunk-wise, which is the architectural guarantee of genuine real-time streaming. One textual inconsistency is worth flagging: Section 4.4 states sampling uses $\tau=1.0$ with $K_{\text{top}}=200$, while Appendix D.2 says temperature 1.0 with no top-k truncation; the two statements about truncation contradict each other and matter for reproduction.
Training scale and data hygiene. AdamW with $\beta_{1}=0.9$, $\beta_{2}=0.95$, weight decay 0.1, gradient clipping at $\ell_{2}$ norm 1.0, learning rate 3e-4, effective batch size 512, trained on 64 NVIDIA H100 GPUs for 30K iterations. Three filters clean the data: a clip-level aspect-ratio and integrity filter drops about 1.5 percent of pairs (landscape, rotated or unreadable files on which HMR2 produces unusable body fits); a frame-level HMR2 validity mask drops about 2.4 percent of frames with is_valid equal to 0, discarding pairs with fewer than 8 valid frames entirely; and about 1.4 percent of pairs are dropped because no roughly 10-second voice clip is available for one participant. The four part-codecs are each trained independently on a single H100 and selected by held-out reconstruction error; during motion-tower training the encoders run once per pair offline and the codes are cached to disk as 16-bit integer bins of shape $(T,22)$.
flowchart LR
subgraph SRC["dyadic input at 12.5 Hz"]
PA["partner audio + partner motion"]
AA["agent audio"]
end
subgraph ST["speech tower: frozen PersonaPlex"]
EMB["17-way interleave
text + Mimi codebooks"]
BLK["32 causal blocks"]
HID["per-layer residual states
H_1 ... H_32"]
SOUT["agent speech stream"]
end
subgraph MT["motion tower: trainable 32 blocks"]
INT["dyadic interleaved stream
[A] m^A_t [B] m^B_t ..."]
SAT["dyadic causal self-attention"]
XAT["cross-attention
time-aligned speech-motion RoPE
causal mask s leq q_pos(t)"]
HEAD["LM head: 18 body codes per frame"]
end
subgraph DEC["part-aware RVQ-VAE decoders"]
PART["4 decoders
upper / hands / lower+trans / face"]
UP["2x temporal upsample
SMPL-X at 25 fps"]
end
PA --> EMB
AA --> EMB
EMB --> BLK
BLK --> HID
BLK --> SOUT
PA --> INT
HID --> XAT
INT --> SAT
SAT --> XAT
XAT --> HEAD
HEAD --> PART
PART --> UP
Reading the structure: the frozen speech tower emits per-layer residual-stream states that serve as cross-attention keys and values for every motion-tower block; the motion tower runs dyadic causal self-attention over the interleaved stream, reads speech through time-aligned RoPE cross-attention, emits 18 body codes per frame from its LM head, and the four part decoders upsample back to 25 fps SMPL-X.
Figure 2: architecture overview. (a) The four part-aware RVQ-VAE decoders (upper, hands, lower plus translation, face, with 6 or 4 quantizers each). (b) The frozen speech tower takes dyadic speech, autoregressively emits agent speech and exposes 32 layers of hidden states; the 32-layer causal motion tower (d_m=1024, h_m=16) operates on the 12.5 fps dyadically interleaved stream, applying dyadic causal self-attention then cross-attention to speech through learned projections (h_c=12, d_c=64) with time-aligned speech-motion RoPE; the LM head outputs motion tokens at 12.5 fps and the decoders upsample 2x to 25 fps SMPL-X pose (paper Figure 2).
Experiments: monadic and dyadic metrics move together
Setup. Training uses roughly 4,000 hours of Seamless Interaction; after filtering, 57,947 pairs (3,435 hours) of dyadic motion remain. Evaluation uses a 330-pair test subset (about 18 hours, both speakers passing an extra audio quality check), keeping only the first 20 seconds of each pair. Baselines are Audio2Photoreal (the diffusion representative) and DualTalk (the transformer representative), both adapted from official source code and retrained, together with the body-only base model, on the same Seamless SMPL body data. Metrics come in two groups: monadic and alignment metrics FGD, Diversity and BeatAlign; dyadic metrics P-FD, the Frechet distance over 132-dimensional per-frame vectors concatenating both speakers' 22x3 joint positions, and Delta-User, the relative P-FD gain when the user motion track is replaced by a shuffled one. Because both baselines consume ground-truth audio of both speakers, DyaPlex is evaluated in a teacher-forced configuration for fairness: ground-truth PersonaPlex hidden states and user motion tokens are provided and only the agent's motion tokens are sampled. All evaluation runs at 25 fps over the 22 SMPL-X body joints (excluding finger and face joints) with root translation set to zero.
| Method | Perceives partner speech | Perceives partner motion | Generates agent speech | Generates agent motion | Dyadic | Causal | Full-duplex | Training data |
|---|---|---|---|---|---|---|---|---|
| SARAH | yes | no (dagger) | no | yes | yes | yes | no | Embody 3D (50 h) |
| Audio2Photoreal | yes | no | no | yes | yes | no | no | custom (8 h) |
| DualTalk | yes | yes | no | yes | yes | no | no | DualTalk (50 h) |
| ViBES | yes | no | yes | yes | no | no | no | Converse 3D (1,000 h) |
| MIBURI | yes | no | yes | yes | no | yes | speech | BEAT2 (70 h) |
| DyaDiT | yes | yes | no | yes | yes | no | no | Seamless (182 h) |
| DyaPlex (ours) | yes | yes | yes | yes | yes | yes | speech and motion | Seamless (4,000 h) |
Table 1: capability comparison across dyadic interaction systems (paper Table 1). The dagger marks SARAH, which perceives the user's 2D floor position but not gestures. DyaPlex is the only method whose motion pathway is full-duplex, and it trains on a corpus 20 to 500 times larger than prior work.
| Method | FGD lower-better (x1e-3) | Diversity closer-to-GT (GT 0.633) | BeatAlign closer-to-GT (GT 0.049) | P-FD lower-better (x1e-3) | Delta-User higher-better |
|---|---|---|---|---|---|
| GT | - | 0.633 | 0.049 | - | - |
| GT (Random) | 13 | 0.683 | 0.050 | 33 | 0% (dagger) |
| Audio2Photoreal | 57 | 0.395 | 0.051 | 72 | 0% (dagger) |
| DualTalk | 161 | 0.305 | - | 163 | +0.3% |
| w/o Self-Attn | 41 | 0.416 | 0.132 | 45 | 0% (dagger) |
| w/o Partner | 39 | 0.725 | 0.064 | 41 | 0% (dagger) |
| w/o Cross-Attn | 41 | 0.708 | 0.080 | 44 | +15% |
| w/o Cross-RoPE | 8.4 | 0.582 | 0.064 | 10 | +31% |
| Ours (body-only) | 5.6 | 0.611 | 0.059 | 7.3 | +31% |
Table 2: quantitative results on the Seamless test set (paper Table 2). FGD and P-FD are lower-is-better; Diversity and BeatAlign are closer-to-GT-better; a positive Delta-User means the generated motion changes when user motion is shuffled at inference. The dagger marks methods with no partner-motion pathway, which score 0 percent by design. DualTalk's BeatAlign is undefined because its collapsed output has no detectable motion.
Reading the main results. On FGD, DyaPlex scores 5.6 against 57 for A2P and 161 for DualTalk, an order-of-margin gap. DualTalk's failure deserves its own paragraph: trained on Seamless body data it collapses to a near-constant body pose, with per-frame body-velocity standard deviation below 1e-5 versus about 1e-2 for DyaPlex, and its velocity term keeps decreasing as the output converges to a static pose. The authors attribute this to two factors: under a mean-squared-error objective the sparse, low-amplitude gestures of conversational body motion make the constant mean-pose solution a strong local optimum, and DualTalk's architecture and loss were designed for audio-to-face lip-sync, where the mapping is near-deterministic and temporally dense, whereas audio-to-body gesture is only weakly and non-deterministically coupled to speech and offers little reliable per-frame signal. On Diversity, DyaPlex's 0.611 is closest to the GT value of 0.633 while A2P (0.395) and DualTalk (0.305) both over-collapse; BeatAlign 0.059 is second best after GT 0.049 and A2P 0.051. On the dyadic side, P-FD is 7.3 against 72 for A2P and 163 for DualTalk, and GT (Random) at 33 confirms P-FD really measures the realism of the joint two-person distribution, since randomly paired ground truth is still closer to matched ground truth than any single-person generation. Delta-User of +31 percent means shuffling the user motion worsens P-FD by 31 percent relatively, i.e. the model genuinely uses the partner's motion; baselines without a partner-motion pathway score 0 percent by definition.
Figure 3: a turn-taking story-reading scene. The pair uses body language to communicate turn taking; the full model produces gestures coherent with the partner and the ground-truth agent, while the variant without partner perception cannot see the partner's gestures and fails to respond coherently (paper Figure 4).
The ablations assign credit cleanly. Removing self-attention collapses everything (FGD 41, P-FD 45, BeatAlign 0.132): causal self-attention over the interleaved sequence is the only carrier of the three dependencies (intra-frame code coherence, within-frame cross-speaker reaction, cross-frame long-range dynamics). Removing partner perception (w/o Partner) raises FGD from 5.6 to 39, about 7x, and P-FD from 7.3 to 41, about 5.6x, while Diversity overshoots to 0.725; this is precisely the difference between dyadic interaction and two stacked monadic generations, and Figure 3 visualizes it on the turn-taking example. Removing cross-attention drops BeatAlign to 0.080 and Delta-User to +15 percent: with no speech context at all, synchronization and conditioning both suffer. Removing only the RoPE inside cross-attention (w/o Cross-RoPE) leaves FGD 8.4 and P-FD 10 respectable but degrades BeatAlign to 0.064, exactly the temporal-alignment precision that RoPE is responsible for, consistent with the diagonal attention maps in Appendix B.
| Comparison | Participants preferring DyaPlex |
|---|---|
| vs. Ground Truth | 29.4% |
| vs. Audio2Photoreal | 66.3% |
| vs. DualTalk | 97.5% |
Table 3: user study with 32 participants (paper Table 3). Participants wore stereo headphones (partner audio in the left channel, agent audio in the right), played the two mesh-rendered videos independently on a web interface, toggled freely between them, and were asked to judge only the agent's body motion.
User study and runtime. Thirty-two participants compared paired mesh-rendered videos on a web interface: identical ground-truth conversation audio and identical ground-truth partner, with only the agent's body motion differing; comparison type (against GT, A2P or DualTalk) and display order were randomized, and participants were instructed to focus on body motion and ignore facial expressions and audio. DyaPlex is preferred over DualTalk by 97.5 percent and over A2P by 66.3 percent, and even against ground truth it takes a 29.4 percent preference rate. On runtime, a single RTX A6000 Ada at 12.5 Hz runs the speech tower at 30 ms per frame and the RVQ-VAE decoder at 0.8 ms per frame; the autoregressive motion tower is the bottleneck at 173 ms per frame with the full 4096-token context, and shrinking the motion context to 1024 tokens (1.8 s) while keeping the full 7.1 s speech context brings it to 80 ms per frame, achieving real-time inference.
Figure 4: cross-attention maps. (a) Without time-aligned speech-motion RoPE the motion tokens cannot cleanly attend to same-frame speech features; (b) with RoPE the attention lights up along the diagonal, showing motion tokens attending to the speech feature of their own frame (paper Figure 5, Appendix B).
| Part | Input dim | Joints / fields | Quantizers | Codebook band |
|---|---|---|---|---|
| Upper | 78 | 13 joints, 6D rotation | K_b=6 | [0, 1024) |
| Hands | 180 | 30 joints, 6D rotation | K_b=6 | [1024, 2048) |
| Lower + translation | 57 | 9 joints 6D rotation + 3D translation velocity | K_b=6 | [2048, 3072) |
| Face | 56 | 6D jaw rotation + 50D FLAME expression | K_f=4 | [3072, 4096) |
Table 4: input decomposition and codebook bands of the four part-aware RVQ-VAEs (paper Appendix D.3 table). The four codecs are independent at codec level and are unified only at the motion-tower input, where their token ids are interleaved into the dyadic stream; the shared 4096-entry vocabulary is the union of the four disjoint 1024-entry bands.
Limitations
One-token-at-a-time decoding of 22 codes (author-stated). The model currently decodes one body token at a time for the 22 codes per frame, which the authors concede may not be optimal; for real robot deployment they suggest exploring chunk-by-chunk decoding (for example six upper-body tokens in one go) or a decomposed depth-transformer design similar to Moshi to speed up inference, and note that other generative backbones such as diffusion models are worth exploring. The model also handles only two users, and extending it to polyadic interaction is named as future work.
Deepfake risk (author-stated). DyaPlex outputs generic SMPL-X body pose parameters rather than photorealistic pixels, so it does not directly produce identity-bearing video; however generated motion could be paired with downstream identity-conditioned video diffusion models to produce realistic-looking videos of fabricated dyadic interactions, creating a risk of accessible deepfakes. The mitigations the authors list are structural: the model is identity-agnostic and conditions on no subject identifier, so impersonating a specific person requires a separate identity-conditioned avatar or voice pipeline with its own safeguards; pixel-footprint-based detectors for AI-generated video exist; and the speech tower is frozen in its public release configuration with no fine-tuning targeting specific real-world voices.
Evaluation protocol and annotation noise (our reading). The main table is produced under a teacher-forced configuration, ground-truth PersonaPlex hidden states plus ground-truth user motion tokens with only agent motion sampled, which is still a distance away from a truly online full-duplex closed loop; error accumulation under streaming deployment is not covered by these numbers. Delta-User shows the model uses partner motion, not that it uses it in a socially sensible direction or magnitude. The body-only base model carries no face, and facial capability appears only in qualitative demos. Finally, Seamless body annotations come from monocular regressors (HMR2 and HaMeR); filtering removes only obvious failures (1.5 percent of pairs, 2.4 percent of frames), so residual regression noise contaminates both the "ground truth" and the training targets, and absolute FGD and P-FD values should be read with that in mind.
Conclusion and outlook
DyaPlex's answer is not a bigger unified backbone but a division of labor: a frozen speech tower keeps zero-shot conversational reasoning and streaming speech, while a trainable motion tower hangs off it as a bypass, drinking per-layer hidden states through cross-attention without touching the prior. The dyadically interleaved sequence gives both speakers' motions a single shared autoregressive prior for the first time, and time-aligned RoPE reduces cross-modal temporal alignment from an implicit mapping that must be learned to an inductive bias written into the structure, whose ideal solution is simply a diagonal. With 4,000 hours of Seamless data the model learns cross-speaker high-order dependencies that small corpora cannot teach, and it ends up setting new state of the art on both monadic and dyadic benchmarks while keeping the strict causality that buys real-time streaming.
The next steps on this route are clear: on the decoding side, chunked or depth-transformer sampling to remove the 22 sequential samples per frame; on the interaction side, extending from dyads to polyadic groups; on the application side, plugging the SMPL-X output into identity-conditioned rendering or robot control to build social robots and virtual humans that genuinely listen, watch and respond. The secondary line, synthetic dyadic interaction data, may itself become the training corpus for the next generation of embodied agents.
In natural dyadic interaction, listening and expressing are never discrete alternating turns, but rather a highly synchronized, streaming process encompassing both speech and physical movements.