arXiv:2505.23525cs.CV2025-05被引 17

用人类偏好优化提升口型同步与表情自然度的高保真动态肖像生成

Hallo4: High-Fidelity Dynamic Portrait Animation via Direct Preference Optimization

  • 基于人类偏好数据直接优化生成结果,提升动画自然度
  • 通过时序运动调制保留高频动作细节,改善口型同步与肢体连贯性
  • 兼容现有扩散模型架构,适合影视动画与虚拟人开发

由于需要精确的口型同步、自然的表情变化和高保真的身体动作动态,基于音频和骨骼运动驱动的高动态、逼真肖像动画生成仍具挑战。我们提出一种以人为中心的偏好对齐扩散框架,包含两项关键创新:首先,引入针对人像动画的直接偏好优化,利用人工标注的偏好数据集,使生成结果更符合视觉感知指标中的运动-视频对齐度与表情自然度;其次,提出的时序运动调制机制通过时间通道重分配与比例特征扩展,解决时空分辨率不匹配问题,将运动条件重塑为维度对齐的潜在特征,有效保留了扩散合成中高频运动细节的保真度。该方法与现有的UNet和DiT-based肖像扩散方法具有互补性。实验表明,在口型-音频同步、表情生动性、身体动作连贯性方面均显著优于基线方法,并在人类偏好评估中取得明显提升。模型与源代码见:https://github.com/fudan-generative-vision/hallo4。

原文摘要 · Abstract (English)

Generating highly dynamic and photorealistic portrait animations driven by audio and skeletal motion remains challenging due to the need for precise lip synchronization, natural facial expressions, and high-fidelity body motion dynamics. We propose a human-preference-aligned diffusion framework that addresses these challenges through two key innovations. First, we introduce direct preference optimization tailored for human-centric animation, leveraging a curated dataset of human preferences to align generated outputs with perceptual metrics for portrait motion-video alignment and naturalness of expression. Second, the proposed temporal motion modulation resolves spatiotemporal resolution mismatches by reshaping motion conditions into dimensionally aligned latent features through temporal channel redistribution and proportional feature expansion, preserving the fidelity of high-frequency motion details in diffusion-based synthesis. The proposed mechanism is complementary to existing UNet and DiT-based portrait diffusion approaches, and experiments demonstrate obvious improvements in lip-audio synchronization, expression vividness, body motion coherence over baseline methods, alongside notable gains in human preference metrics. Our model and source code can be found at: https://github.com/fudan-generative-vision/hallo4.

肖像动画扩散模型口型同步偏好优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。