arXiv:2606.27779cs.CV2026-06中稿 · ECCV

用双路径模型让对话面部动画更懂情绪、更自然。

MindFlow: Harmonizing Cognitive Semantics and Acoustic Dynamics for Facial Animation Generation in Dyadic Conversations

论文配图:MindFlow: Harmonizing Cognitive Semantics and Acoustic Dynamics for Facial Animation Generation in Dyadic Conversations
图 1 · 摘自论文原文
  • 分两条路:一条理解情绪变化,一条精准控制表情动作
  • 比现有方法在语义贴合度和动作自然度上均提升明显
  • 适合做虚拟人对话、影视动画的开发者

生成双人对话中的逼真面部动画,需兼顾高层认知意图与底层运动反射。现有方法在对话语境语义理解与精细动态控制方面仍有不足。本文提出 MindFlow,受神经科学中腹侧-背侧通路启发,构建双路径生成框架,实现深层语义推理与细粒度控制的协同。腹侧模块将传统句-动作模式改为新提出的块-状态模式,将原始声学信号建模为上下文感知、持续演化的情绪状态链,捕捉细微的副语言特征及句中情绪变化。背侧模块采用条件自回归流匹配网络,以高频声学线索驱动高保真面部运动,并通过选择性声学注入器实现自适应音频门控,确保说话与倾听动态下的鲁棒性,避免干扰。大量实验表明,相比最先进基线,MindFlow 在语义恰当性和动作自然性上均有显著提升。

原文摘要 · Abstract (English)

Generating lifelike facial animation for dyadic conversations requires reconciling high-level cognitive intent with precise low-level motor reflexes, yet existing methods fall short in the semantic understanding of dialogue context and in precise dynamic control. In this paper, we propose MindFlow, a dual-pathway generative framework inspired by the Ventral-Dorsal pathway model in neuroscience, which decouples generation into two collaborative streams, thereby harmonizing deep semantic reasoning with fine-grained control. In the Ventral module, we transform the conventional Sentence-Action approach into a novel Chunk-State approach that models raw acoustic streams as a context-aware, evolving emotional state chain, capturing subtle paralinguistic nuances and mid-utterance emotional shifts missed by sentence-level modeling. The Dorsal module features a conditional autoregressive flow matching network for high-fidelity facial motion, driven by high-frequency acoustic cues and modulated by emotion states, plus a Selective Acoustic Injector for adaptive audio gating to ensure robustness in talking-and-listening dynamics without interference. Extensive experiments demonstrate that MindFlow achieves superior semantic appropriateness and motion naturalness compared to state-of-the-art baselines.

面部动画对话生成情绪建模双路径

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。