让虚拟头像更自然,能根据上下文和情绪生成恰当表情。
ECHO: Towards Emotionally Appropriate and Contextually Aware Interactive Head Generation
- 用长时上下文理解模块增强表情的语境合理性
- 分块解耦注意力机制提升口型同步与面部细节质量
- 适合需要真实互动感的虚拟人/元宇宙场景
在自然面对面交流中,参与者会无缝切换发言与倾听,产生的面部行为(FBs)受长期上下文影响,具备情境适切性和情绪合理性。交互式头像生成(IHG)旨在合成能模拟此类能力的逼真虚拟头像视频。现有方法通常依赖短期窗口内的双路信号(用户行为与预设音频)驱动头像口型对齐与非语言面部行为生成,但存在两大问题:(i) 仅依赖短时行为线索,缺乏长程上下文建模,导致生成行为缺乏情境适切性;(ii) 双路信号纠缠融合引发跨信号干扰,可能损害说话时口部同步。为此,我们提出ECHO框架,包含两个核心组件:长程上下文理解(LCU)模块,用于融合行为驱动动态与语言驱动情感语义,提升生成面部行为的情境适切性与情绪合理性;以及分块空间感知解耦交叉注意力调制(SDCM)模块,保持自音频驱动口型的同时,自适应整合用户上下文行为线索用于非口部区域,并配合两阶段训练范式,共同提升口型同步与视觉保真度。大量实验验证了各组件有效性及ECHO在IHG任务中的优越性能。
原文摘要 · Abstract (English)
In natural face-to-face interaction, participants seamlessly alternate between speaking and listening, producing facial behaviors (FBs) that are finely informed by long-range context and naturally exhibit contextual appropriateness and emotional rationality. Interactive Head Generation (IHG) aims to synthesize lifelike avatar head video emulating such capabilities. Existing IHG methods typically condition on dual-track signals (i.e., human user's behaviors and pre-defined audio for avatar) within a short temporal window, jointly driving generation of avatar's audio-aligned lip articulation and non-verbal FBs. However, two main challenges persist in these methods: (i) the reliance on short-clip behavioral cues without long-range contextual modeling leads them to produce facial behaviors lacking contextual appropriateness; and (ii) the entangled, role-agnostic fusion of dual-track signals empirically introduces cross-signal interference, potentially compromising lip-region synchronization during speaking. To this end, we propose ECHO, a novel IHG framework comprising two key components: a Long-range Contextual Understanding (LCU) component that facilitates contextual understanding of both behavior-grounded dynamics and linguistic-driven affective semantics to promote contextual appropriateness and emotional rationality of synthesized avatar FBs; and a block-wise Spatial-aware Decoupled Cross-attention Modulation (SDCM) module, that preserves self-audio-driven lip articulation while adaptively integrating user contextual behavioral cues for non-lip facial regions, complemented by our designed two-stage training paradigm, to jointly enhance lip synchronization and visual fidelity. Extensive experiments demonstrate the effectiveness of proposed components and ECHO's superior IHG performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。