arXiv:2508.10566cs.CV2025-08

融合音视频特征,让虚拟人脸说话更自然真实。

HM-Talker: Hybrid Motion Modeling for High-Fidelity Talking Head Synthesis

  • 用音视频联合建模提取运动线索,兼顾身份特性和语音驱动。
  • 在多个数据集上实现更高画质和唇动同步精度。
  • 适合需要高保真人脸动画的影视、虚拟主播场景。

音频驱动的人脸生成面临个性化与泛化之间的根本权衡,限制了实际应用。隐式模型虽具泛化性但常导致头部动作不稳和唇形不同步;显式方法虽引入3DMM或动作单元(AUs)等先验知识,却易产生表情呆板或泛化能力不足。为此,本文提出HM-Talker,通过协同融合显式发音线索与隐式语调特征,刻画身份特异性动态并支持音频驱动的泛化。其核心包括:i)跨模态映射模块(CMMM),从音视频中提取全面的运动特征词汇;ii)混合运动建模模块(HMMM),采用随机特征配对(SFP)策略动态融合隐式与显式特征进行运动合成。该设计实现下颌区域运动的迭代优化,在身份特定与仅音频驱动的目标间交替提升。大量实验表明,HM-Talker在多种设置下均优于现有最优方法,在视觉真实感与唇形同步精度上均有显著提升。

原文摘要 · Abstract (English)

Audio-driven talking head generation faces a fundamental trade-off between personalization and generalization, limiting its practical application. Implicit models often achieve generalization at the cost of structural incoherence, resulting in unstable head motion and inaccurate lip synchronization. While explicit methods incorporate geometric and anatomical priors such as 3D Morphable Models (3DMMs), which parameterize facial geometry, or Action Units (AUs), which code facial muscle movements--they tend to produce overly neutral expressions or suffer from limited generalization. To resolve this conflict, we present HM-Talker, an audio-driven talking head framework that synergistically integrates explicit articulatory cues with implicit prosodic features to characterize identity-specific dynamics while enabling audio-driven generalization. Its distinctive features can be summarized as: i) the Cross-Modal Mapping Module (CMMM) that extracts a comprehensive vocabulary of motion cues from audio and video, and ii) the Hybrid Motion Modeling Module (HMMM) that employs a Stochastic Feature Pairing (SFP) strategy to dynamically merge paired implicit and explicit features for motion synthesis. This design facilitates an iterative optimization of the lower face motion, alternating between identity-specific and identity-agnostic (audio-only) objectives. Extensive experiments demonstrate that HM-Talker outperforms state-of-the-art methods in both visual realism and lip-sync accuracy across diverse settings.

人脸生成语音驱动运动建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。