arXiv:2507.00472cs.CV2025-07ICCV被引 15

实时对话中生成自然交互头动作,提升虚拟人表现力

ARIG: Autoregressive Interactive Head Generation for Real-time Conversations

  • 采用自回归框架逐帧生成,避免片段式处理的延迟问题
  • 通过扩散过程建模连续动作分布,提升运动预测精度
  • 融合双向行为理解与对话状态识别,实现更真实的互动反应

面对面交流是常见的人类活动,推动了交互式头部动作生成的研究。虚拟代理可根据对方或自身的音频或动作信号,具备听和说的能力。然而,以往基于片段的生成范式或显式切换监听/发言机制的方法,在未来信号获取、上下文行为理解及切换平滑性方面存在局限,难以实现实时且逼真的交互。本文提出一种基于自回归(AR)的逐帧生成框架ARIG,实现更真实的实时生成。为实现实时性,将动作预测建模为非向量量化自回归过程;不同于离散码本索引预测,采用扩散过程表示动作分布,提升连续空间中的预测准确性。为增强交互真实性,强调交互行为理解(IBU)与对话状态理解(CSU)。在IBU中,基于双路径双模态信号,通过双向融合学习捕捉短程行为,并进行长程上下文理解;在CSU中,利用语音活动信号与IBU上下文特征,理解实际对话中的多种状态(如打断、反馈、停顿等),作为最终渐进式动作预测的条件。大量实验验证了模型的有效性。

原文摘要 · Abstract (English)

Face-to-face communication, as a common human activity, motivates the research on interactive head generation. A virtual agent can generate motion responses with both listening and speaking capabilities based on the audio or motion signals of the other user and itself. However, previous clip-wise generation paradigm or explicit listener/speaker generator-switching methods have limitations in future signal acquisition, contextual behavioral understanding, and switching smoothness, making it challenging to be real-time and realistic. In this paper, we propose an autoregressive (AR) based frame-wise framework called ARIG to realize the real-time generation with better interaction realism. To achieve real-time generation, we model motion prediction as a non-vector-quantized AR process. Unlike discrete codebook-index prediction, we represent motion distribution using diffusion procedure, achieving more accurate predictions in continuous space. To improve interaction realism, we emphasize interactive behavior understanding (IBU) and detailed conversational state understanding (CSU). In IBU, based on dual-track dual-modal signals, we summarize short-range behaviors through bidirectional-integrated learning and perform contextual understanding over long ranges. In CSU, we use voice activity signals and context features of IBU to understand the various states (interruption, feedback, pause, etc.) that exist in actual conversations. These serve as conditions for the final progressive motion prediction. Extensive experiments have verified the effectiveness of our model.

交互生成自回归虚拟人

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。