arXiv:2512.15340cs.CV2025-12被引 4

建模对话中头动的因果关系,让虚拟人互动更自然

Towards Seamless Interaction: Causal Turn-Level Modeling of Interactive 3D Conversational Head Dynamics

  • 用因果注意力建模每轮对话的视听信息流
  • 在DualTalk上降低15-30%的误差,跨数据集表现稳定
  • 适合做虚拟人、交互机器人动画生成的研究者

人类对话包含语音与头部点头、视线转移、面部表情等非语言线索的持续交互,这些线索传递注意力与情绪。构建有表现力的虚拟人和交互机器人需精准建模三维对话头动的双向动态。现有方法常将说话与倾听视为独立过程,或依赖非因果的全序列建模,导致轮次间时间连贯性差。本文提出TIMAR(Turn-level Interleaved Masked AutoRegression)框架,以因果方式建模3D对话头动,将对话视为交错的音视频上下文。它在每轮内融合多模态信息,并通过轮级因果注意力累积对话历史;轻量级扩散头生成连续3D头动,捕捉协调性与表达多样性。在DualTalk基准测试中,TIMAR使测试集的Fréchet距离与均方误差降低15-30%,在分布外数据上也取得相似提升。源代码已公开于https://github.com/CoderChen01/towards-seamless-interaction。

原文摘要 · Abstract (English)

Human conversation involves continuous exchanges of speech and nonverbal cues such as head nods, gaze shifts, and facial expressions that convey attention and emotion. Modeling these bidirectional dynamics in 3D is essential for building expressive avatars and interactive robots. However, existing frameworks often treat talking and listening as independent processes or rely on non-causal full-sequence modeling, hindering temporal coherence across turns. We present TIMAR (Turn-level Interleaved Masked AutoRegression), a causal framework for 3D conversational head generation that models dialogue as interleaved audio-visual contexts. It fuses multimodal information within each turn and applies turn-level causal attention to accumulate conversational history, while a lightweight diffusion head predicts continuous 3D head dynamics that captures both coordination and expressive variability. Experiments on the DualTalk benchmark show that TIMAR reduces Fréchet Distance and MSE by 15-30% on the test set, and achieves similar gains on out-of-distribution data. The source code has been released at https://github.com/CoderChen01/towards-seamless-interaction.

3D生成对话建模扩散模型虚拟人

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。