让3D头像表情可独立控制,实现自然情感表达。
Giving Faces Their Feelings Back: Explicit Emotion Control for Feedforward Single-Image 3D Head Avatars

- 用双路调制机制将情绪作为独立信号注入现有模型
- 支持跨身份情感迁移与平滑情绪插值,保持还原精度
- 适合需要精细情感控制的虚拟人、影视特效等场景
我们提出一种前馈式单图3D头像重建框架,实现显式情绪控制。不同于以往将情绪隐含于几何或外观中的方法,我们将情绪视为可独立操作的第一类控制信号。通过双路径调制机制,在不修改原有架构的前提下,将情绪注入现有模型:几何调制在原始参数空间进行情绪条件归一化,解耦情绪状态与语音驱动的嘴型变化;外观调制捕捉超越几何的身份感知、情绪依赖视觉特征。为支持该设定下的学习,我们构建了一个时间同步、情绪一致的多身份数据集,通过跨身份转移对齐的情感动态生成。集成到多个顶尖骨干网络后,本框架在保持重建与重演保真度的同时,实现了可控情感迁移、解耦操作与平滑情绪插值,推动了更具表现力与可扩展性的3D头像发展。
原文摘要 · Abstract (English)
We present a framework for explicit emotion control in feed-forward, single-image 3D head avatar reconstruction. Unlike existing pipelines where emotion is implicitly entangled with geometry or appearance, we treat emotion as a first-class control signal that can be manipulated independently and consistently across identities. Our method injects emotion into existing feed-forward architectures via a dual-path modulation mechanism without modifying their core design. Geometry modulation performs emotion-conditioned normalization in the original parametric space, disentangling emotional state from speech-driven articulation, while appearance modulation captures identity-aware, emotion-dependent visual cues beyond geometry. To enable learning under this setting, we construct a time-synchronized, emotion-consistent multi-identity dataset by transferring aligned emotional dynamics across identities. Integrated into multiple state-of-the-art backbones, our framework preserves reconstruction and reenactment fidelity while enabling controllable emotion transfer, disentangled manipulation, and smooth emotion interpolation, advancing expressive and scalable 3D head avatars.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。