arXiv:2510.04712cs.CVcs.HC2025-10中稿 · ACM Multimedia被引 5

让对话表情生成更自然多样,模拟真人反应的时序规律。

ReactDiff: Fundamental Multiple Appropriate Facial Reaction Diffusion Model

  • 用时序扩散模型生成表情,融入面部运动与动作单元约束
  • 在REACT2024数据集上表现领先,兼顾真实感、多样性与情境适配性
  • 适合需要逼真人物交互的虚拟助手、游戏角色开发

在双人对话中自动生成多样且类人化的面部反应仍是人机交互系统的关键挑战。现有方法难以建模真实人类反应固有的随机性与动态特性。为此,我们提出ReactDiff,一种新型时序扩散框架,用于生成对任意对话上下文均恰当的多样化面部反应。核心思想是:可信的人类反应具有时间上的平滑性与连贯性,并受人类面部解剖结构约束。ReactDiff在扩散过程中引入两项关键先验:(i) 时序面部行为动力学,(ii) 面部动作单元依赖关系。这两项约束引导模型走向真实的人类反应流形,避免视觉上不自然的抖动、不稳定过渡、异常表情等伪影。在REACT2024数据集上的大量实验表明,该方法不仅在反应质量上达到当前最优,还在多样性与情境适配性方面表现优异。

原文摘要 · Abstract (English)

The automatic generation of diverse and human-like facial reactions in dyadic dialogue remains a critical challenge for human-computer interaction systems. Existing methods fail to model the stochasticity and dynamics inherent in real human reactions. To address this, we propose ReactDiff, a novel temporal diffusion framework for generating diverse facial reactions that are appropriate for responding to any given dialogue context. Our key insight is that plausible human reactions demonstrate smoothness, and coherence over time, and conform to constraints imposed by human facial anatomy. To achieve this, ReactDiff incorporates two vital priors (spatio-temporal facial kinematics) into the diffusion process: i) temporal facial behavioral kinematics and ii) facial action unit dependencies. These two constraints guide the model toward realistic human reaction manifolds, avoiding visually unrealistic jitters, unstable transitions, unnatural expressions, and other artifacts. Extensive experiments on the REACT2024 dataset demonstrate that our approach not only achieves state-of-the-art reaction quality but also excels in diversity and reaction appropriateness.

面部生成扩散模型对话系统

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。