让3D说话头的微表情更自然可控,提升真实感
SubtleTalk: Generating Controllable Weakly-correlated Facial Dynamics for 3D Talking Heads via Residual Flow Matching

- 用语音+情感信号多条件控制微表情生成
- 通过残差流匹配实现多样化的非确定性动作
- 自建3900人规模数据集,提升上半脸追踪精度
基于语音的3D面部动画旨在从语音中合成逼真且时间连贯的面部运动。尽管在口型同步方面取得显著进展,但眉毛动作、眨眼和头部运动等弱相关动态仍难以真实建模,常表现为静态或不自然重复。我们归因于三点:(a) 弱相关动态条件不足;(b) 确定性回归表达能力有限;(c) 数据瓶颈来自不可靠的上半脸伪标签与数据集多样性不足。为此,我们提出SubtleTalk框架,通过多条件建模与残差流匹配生成自然可控的弱相关面部动态。首先,引入可解释控制信号(语调、区域强度、情绪值-唤醒度),显式捕捉微表情的时间、幅度与情感变化。其次,基于稳定语音驱动的运动先验构建残差流匹配,使模型能捕捉超出确定性预测的随机偏差。第三,构建大型3D面部动画数据集SubtleTalk-Face,包含约3,900个身份和74小时数据,采用简单可扩展的伪标注流程,提升上半脸追踪精度与帧级VA标注质量。大量实验表明,该方法显著提升弱相关动态的真实感与多样性,同时保持精确口型同步。
原文摘要 · Abstract (English)
Audio-driven 3D facial animation aims to synthesize realistic and temporally coherent motions from speech. Despite notable progress in lip synchronization, weakly correlated dynamics, including eyebrow movements, eye blinks, and head motion, which are essential to photorealistic facial animation, remain difficult to model faithfully and often appear static or unnaturally repetitive. We attribute this limitation to three factors: (a) insufficient conditioning for weakly correlated dynamics; (b) the limited ability of deterministic regression to capture diverse motion patterns; (c) data bottlenecks from unreliable upper-face pseudo-labels and limited dataset diversity. To address these issues, we propose SubtleTalk, a framework for generating natural and controllable weakly correlated facial dynamics via multi-condition modeling and residual flow matching. First, to compensate for the limited guidance of speech alone, we introduce interpretable controls, including prosody, regional intensity, and Valence-Arousal signals, to explicitly capture the timing, magnitude, and affective variation of weakly correlated dynamics. Second, to overcome the limited expressiveness of deterministic regression, we build residual flow matching based on a stable speech-driven motion prior, allowing the model to capture stochastic deviations beyond deterministic prediction. Third, to alleviate the data bottleneck, we construct SubtleTalk-Face, a large-scale 3D facial animation dataset comprising about 3,900 identities and 74 hours of data, built via a simple and scalable pseudo-labeling pipeline and featuring improved upper-face tracking and frame-level VA annotations. Extensive experiments demonstrate that our method significantly improves the realism and diversity of weakly correlated facial dynamics while preserving accurate lip synchronization.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。