arXiv:2508.11255cs.CV2025-08AAAI被引 4

用多维度偏好优化让人脸动画更自然、对口型更准。

FantasyTalking2: Timestep-Layer Adaptive Preference Optimization for Audio-Driven Portrait Animation

  • 构建可量化多维度偏好的奖励模型Talking-Critic
  • 在41万组对比数据上训练,提升对口型与动作自然度
  • 分层时间自适应融合专家模块,避免各目标相互干扰

近期音频驱动人脸动画进展显著,但现有方法难以同时满足运动自然性、对口型准确性和视觉质量等多维人类偏好。这源于优化冲突目标的困难以及高质量多维偏好标注数据稀缺。为此,我们提出Talking-Critic,一种多模态奖励模型,可学习符合人类偏好的奖励函数以量化生成视频对多维期望的满足程度。基于该模型,我们构建了包含41万组偏好对的大规模多维人类偏好数据集Talking-NSQ。最后,我们提出时序-层级自适应多专家偏好优化(TLPO)框架,将偏好解耦为专用专家模块,在时间步与网络层间融合,实现无相互干扰的全维度精细优化。实验表明,Talking-Critic显著优于现有方法在人类偏好匹配上的表现;而TLPO在对口型准确率、运动自然性与视觉质量上均显著超越基线模型,定性与定量评估均表现优异。

原文摘要 · Abstract (English)

Recent advances in audio-driven portrait animation have demonstrated impressive capabilities. However, existing methods struggle to align with fine-grained human preferences across multiple dimensions, such as motion naturalness, lip-sync accuracy, and visual quality. This is due to the difficulty of optimizing among competing preference objectives, which often conflict with one another, and the scarcity of large-scale, high-quality datasets with multidimensional preference annotations. To address these, we first introduce Talking-Critic, a multimodal reward model that learns human-aligned reward functions to quantify how well generated videos satisfy multidimensional expectations. Leveraging this model, we curate Talking-NSQ, a large-scale multidimensional human preference dataset containing 410K preference pairs. Finally, we propose Timestep-Layer adaptive multi-expert Preference Optimization (TLPO), a novel framework for aligning diffusion-based portrait animation models with fine-grained, multidimensional preferences. TLPO decouples preferences into specialized expert modules, which are then fused across timesteps and network layers, enabling comprehensive, fine-grained enhancement across all dimensions without mutual interference. Experiments demonstrate that Talking-Critic significantly outperforms existing methods in aligning with human preference ratings. Meanwhile, TLPO achieves substantial improvements over baseline models in lip-sync accuracy, motion naturalness, and visual quality, exhibiting superior performance in both qualitative and quantitative evaluations. Ours project page: https://fantasy-amap.github.io/fantasy-talking2/

人脸动画偏好优化扩散模型对口型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。