arXiv:2505.07901cs.LGcs.AI2025-05被引 4

用扩散模型生成更自然的对话回应表情,提升人机交互真实感。

Latent Behavior Diffusion for Sequential Reaction Generation in Dyadic Setting

  • 通过自编码器压缩输入行为特征,生成紧凑的潜在表示。
  • 在潜在空间中用扩散模型非自回归生成多样且贴合语境的表情反应。
  • 相比现有方法,在对话反应生成任务中表现更优,适合虚拟助手研发。

二人对话中的反应生成任务旨在合成与对话伙伴行为高度一致的面部反应,以增强类人交互模拟的真实性和有效性。本文提出一种新方法——潜变量行为扩散模型,包含上下文感知自编码器和基于扩散的条件生成器,解决从输入说话者行为生成多样化且语境相关的面部反应的难题。自编码器将高维输入特征压缩,捕捉听者反应中的动态模式,将复杂数据浓缩为简洁的潜在表示,从而促进更具表现力和语境恰当的反应合成。扩散条件生成器在自编码器生成的潜在空间上运行,以非自回归方式预测真实感强的面部反应,可生成反映细微对话线索和情绪状态差异的多样性反应。实验结果表明,该方法在二人反应生成任务中优于现有方法。

原文摘要 · Abstract (English)

The dyadic reaction generation task involves synthesizing responsive facial reactions that align closely with the behaviors of a conversational partner, enhancing the naturalness and effectiveness of human-like interaction simulations. This paper introduces a novel approach, the Latent Behavior Diffusion Model, comprising a context-aware autoencoder and a diffusion-based conditional generator that addresses the challenge of generating diverse and contextually relevant facial reactions from input speaker behaviors. The autoencoder compresses high-dimensional input features, capturing dynamic patterns in listener reactions while condensing complex input data into a concise latent representation, facilitating more expressive and contextually appropriate reaction synthesis. The diffusion-based conditional generator operates on the latent space generated by the autoencoder to predict realistic facial reactions in a non-autoregressive manner. This approach allows for generating diverse facial reactions that reflect subtle variations in conversational cues and emotional states. Experimental results demonstrate the effectiveness of our approach in achieving superior performance in dyadic reaction synthesis tasks compared to existing methods.

表情生成扩散模型对话系统

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。