arXiv:2505.04996cs.GRcs.CV2025-05中稿 · ICMR 2025被引 2

让说话人和听众的肢体互动真实自然,实现双向动态响应。

Inter-Diffusion Generation Model of Speakers and Listeners for Effective Communication

  • 引入听众全身体态,用跨体交互扩散机制建模双向互动。
  • 生成动作更自然连贯,语音与手势同步性显著提升。
  • 适合虚拟人交互、智能客服等需要真实对话体验的场景。

全身动作在自然交互中起关键作用,但现有研究多聚焦说话人动作生成,忽视听众的反馈作用及双方动态互动。本文首次提出说话人与听众的跨扩散生成模型,将听众全身体态纳入生成框架。基于先进扩散模型架构,创新引入交互条件与GAN结构,增大去噪步长,使模型能根据说话人语义动态生成动作,并实时响应听众反馈,实现双向协同交互。大量实验表明,相比当前最优方法,本模型在动作自然度、连贯性及语音-手势同步性上均有显著提升。主观评价显示用户认为生成交互更贴近真实人际沟通;客观指标也全面优于基线方法,为有效通信提供更强支持。

原文摘要 · Abstract (English)

Full-body gestures play a pivotal role in natural interactions and are crucial for achieving effective communication. Nevertheless, most existing studies primarily focus on the gesture generation of speakers, overlooking the vital role of listeners in the interaction process and failing to fully explore the dynamic interaction between them. This paper innovatively proposes an Inter-Diffusion Generation Model of Speakers and Listeners for Effective Communication. For the first time, we integrate the full-body gestures of listeners into the generation framework. By devising a novel inter-diffusion mechanism, this model can accurately capture the complex interaction patterns between speakers and listeners during communication. In the model construction process, based on the advanced diffusion model architecture, we innovatively introduce interaction conditions and the GAN model to increase the denoising step size. As a result, when generating gesture sequences, the model can not only dynamically generate based on the speaker's speech information but also respond in realtime to the listener's feedback, enabling synergistic interaction between the two. Abundant experimental results demonstrate that compared with the current state-of-the-art gesture generation methods, the model we proposed has achieved remarkable improvements in the naturalness, coherence, and speech-gesture synchronization of the generated gestures. In the subjective evaluation experiments, users highly praised the generated interaction scenarios, believing that they are closer to real life human communication situations. Objective index evaluations also show that our model outperforms the baseline methods in multiple key indicators, providing more powerful support for effective communication.

动作生成人机交互扩散模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。