让虚拟头像能精准控制眼神、点头节奏和情绪,实现自然对话互动。
STEER: Steerable Dyadic Head Avatars

- 将对话行为分解为眼神、头部节奏和情绪三类可调参数,实现精细控制。
- 在无标注的野外双人视频中自动提取行为伪标签,构建新数据集。
- 无需重训练即可驱动高保真3D头像,支持实时交互与多维度编辑。
面部动作与表情是面对面交流的核心,传递轮流发言、注意力、同意与参与度等非语言信息。尽管语音驱动的面部动画在唇形同步和音频引导运动生成方面取得进展,但多数方法将对话行为视为音频的副产品,或仅提供粗粒度的情绪序列控制。因此,眼神接触与回避、节奏性头部运动及情绪等关键非语言信号仍难以显式调控。本文提出STEER,一种可调控的3D双人对话头像运动先验模型。它将对话行为显式分解为眼神、头部节奏与情绪三个控制维度,使用户可调节虚拟角色如何倾听、反应与互动。由于公开双人语料库缺乏对齐的行为标注,我们构建了一套追踪与标注流程,从真实场景视频中恢复行为伪标签。基于因果流匹配的Transformer模型,学习在音频、对方动作、情绪及行为控制输入下生成伙伴感知的目标运动。此外,通过扩展通用高斯头像先验,引入从追踪参数映射到驱动空间的可学习映射,实现无需重训练的高保真头像可控动画。STEER在运动质量、动态性与多样性上优于近期基线,与伙伴耦合表现相当,并支持眼神、头部节奏与情绪的联合编辑,且具备实时交互能力。代码与数据标注已开源。
原文摘要 · Abstract (English)
Facial movement and expression are central to face-to-face communication, conveying turn-taking, attention, agreement, and engagement alongside speech. While speech-driven facial animation has made strong progress in lip synchronization and audio-conditioned motion generation, most methods treat conversational behavior as an emergent byproduct of audio, or expose only coarse sequence-level affect control. As a result, key non-verbal channels such as gaze contact and aversion, rhythmic head motion, and emotion remain difficult to explicitly control. We present STEER, a controllable 3D dyadic motion prior for reactive conversational head avatars. STEER factorizes conversational behavior into explicit controls for gaze, head rhythm, and emotion, allowing users to steer how an avatar listens, reacts, and engages with a conversation partner. Since temporally aligned annotations for these behaviors are not available in public dyadic corpora, we introduce a tracking and annotation pipeline that recovers behavioral pseudo-labels from in-the-wild dyadic video. A causal flow-matching transformer then learns partner-aware target motion conditioned on audio, partner motion, emotion and the proposed behavioral controls. We further embed STEER in a photorealistic avatar pipeline by extending a Universal Gaussian Head-Avatar Prior with a learned mapping from tracked parametric motion into its avatar-driving space. This enables controllable animation of high-fidelity Gaussian head avatars without re-training the underlying avatar model. STEER outperforms recent dyadic motion baselines on motion quality, dynamics, and diversity, remains competitive on partner coupling, and enables gaze, head-rhythm, and emotion edits together with an interactive live deployment. We make our code and dataset annotations available at our webpage.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。