arXiv:2505.14151cs.CVcs.MM2025-05中稿 · Neural Networks被引 3

用潜空间扩散模型生成更真实多样的面部反应。

ReactDiff: Latent Diffusion for Facial Reaction Generation

  • 融合多模态注意力与潜空间扩散,实现精细跨模态交互。
  • 在相关性(0.26)和多样性(0.094)上优于现有方法。
  • 适合影视合成、虚拟人交互等需要自然表情生成场景。

给定说话者的音视频片段,面部反应生成旨在预测听者的面部反应。挑战在于捕捉音视频间的相关性,同时平衡适当性、真实性和多样性。以往工作多关注单模态输入或简化反应映射,近期方法如PerFRDiff虽探索了多模态输入和一对多反应映射,但仍存局限。本文提出面部反应扩散框架ReactDiff,首次将多模态Transformer与潜空间条件扩散相结合,通过类内与类间注意力实现细粒度多模态交互,潜空间编码器-解码器间的扩散过程生成多样且上下文合理的结果。实验表明,ReactDiff显著优于现有方法,在面部反应相关性达0.26、多样性得分为0.094的同时保持良好真实性。

原文摘要 · Abstract (English)

Given the audio-visual clip of the speaker, facial reaction generation aims to predict the listener's facial reactions. The challenge lies in capturing the relevance between video and audio while balancing appropriateness, realism, and diversity. While prior works have mostly focused on uni-modal inputs or simplified reaction mappings, recent approaches such as PerFRDiff have explored multi-modal inputs and the one-to-many nature of appropriate reaction mappings. In this work, we propose the Facial Reaction Diffusion (ReactDiff) framework that uniquely integrates a Multi-Modality Transformer with conditional diffusion in the latent space for enhanced reaction generation. Unlike existing methods, ReactDiff leverages intra- and inter-class attention for fine-grained multi-modal interaction, while the latent diffusion process between the encoder and decoder enables diverse yet contextually appropriate outputs. Experimental results demonstrate that ReactDiff significantly outperforms existing approaches, achieving a facial reaction correlation of 0.26 and diversity score of 0.094 while maintaining competitive realism. The code is open-sourced at \href{https://github.com/Hunan-Tiger/ReactDiff}{github}.

面部生成扩散模型多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。