arXiv:2409.19580cs.CV2024-09被引 7

通过区域监督与运动模糊建模,提升人像动画的细节真实感。

High Quality Human Image Animation using Regional Supervision and Motion Blur Condition

论文配图:High Quality Human Image Animation using Regional Supervision and Motion Blur Condition
图 1 · 摘自论文原文
  • 对人脸和手部等关键区域施加局部监督,增强细节还原。
  • 显式建模运动模糊,使动态画面更自然,质量提升21%以上。
  • 适合关注高保真人像生成与视频扩散模型优化的研究者。

近期视频扩散模型在实现具时间连贯性的人像动画方面取得了显著进展。然而,现有方法常忽视面部和手部等关键区域的局部监督,且未显式建模运动模糊,导致合成结果缺乏真实感。为此,本文首先在关键区域引入局部监督以提升人脸与手部的忠实度;其次,显式建模运动模糊以进一步改善外观质量;最后,探索适用于高分辨率人像动画的新训练策略,整体提升生成保真度。实验表明,该方法在HumanDance数据集上相较于最强基线,在重建精度(L1)和感知质量(FVD)上分别提升超过21.0%和57.4%,显著优于现有方法。代码与模型将公开。

原文摘要 · Abstract (English)

Recent advances in video diffusion models have enabled realistic and controllable human image animation with temporal coherence. Although generating reasonable results, existing methods often overlook the need for regional supervision in crucial areas such as the face and hands, and neglect the explicit modeling for motion blur, leading to unrealistic low-quality synthesis. To address these limitations, we first leverage regional supervision for detailed regions to enhance face and hand faithfulness. Second, we model the motion blur explicitly to further improve the appearance quality. Third, we explore novel training strategies for high-resolution human animation to improve the overall fidelity. Experimental results demonstrate that our proposed method outperforms state-of-the-art approaches, achieving significant improvements upon the strongest baseline by more than 21.0% and 57.4% in terms of reconstruction precision (L1) and perceptual quality (FVD) on HumanDance dataset. Code and model will be made available.

人像动画扩散模型运动模糊局部监督

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。