用少样本视频教会模型生成细腻的人体动作动画
Learning to Animate Images from A Few Videos to Portray Delicate Human Actions
- 通过跨视频运动特征迁移学习通用动作模式
- 仅用16个以内视频训练,生成动作自然流畅
- 适合影视制作等需精细动作生成的场景
尽管近期进展显著,视频生成模型在将静态图像转化为展现细腻人体动作的视频方面仍面临挑战,尤其当涉及罕见或新颖动作且训练数据有限时。本文探索仅使用16个或更少视频来学习动画生成的方法,对影视制作等真实应用极具价值。在少样本条件下学习可泛化的运动模式并实现从参考图像平滑过渡,极具挑战性。我们提出FLASH(Few-shot Learning to Animate and Steer Humans),通过利用相同动作但不同外观的另一视频的运动特征和跨帧对应关系进行视频重建,迫使模型学习可迁移的运动模式,缓解有限数据下的过拟合问题。此外,FLASH在解码器中增加额外层,将参考图像细节传递至生成帧,提升过渡平滑度。人工评估显示,488次对比中有65.78%的选择偏好FLASH优于基线方法。强烈建议访问官网观看视频:https://lihaoxin05.github.io/human_action_animation/,因运动伪影仅在视频中明显。
原文摘要 · Abstract (English)
Despite recent progress, video generative models still struggle to animate static images into videos that portray delicate human actions, particularly when handling uncommon or novel actions whose training data are limited. In this paper, we explore the task of learning to animate images to portray delicate human actions using a small number of videos -- 16 or fewer -- which is highly valuable for real-world applications like video and movie production. Learning generalizable motion patterns that smoothly transition from user-provided reference images in a few-shot setting is highly challenging. We propose FLASH (Few-shot Learning to Animate and Steer Humans), which learns generalizable motion patterns by forcing the model to reconstruct a video using the motion features and cross-frame correspondences of another video with the same motion but different appearance. This encourages transferable motion learning and mitigates overfitting to limited training data. Additionally, FLASH extends the decoder with additional layers to propagate details from the reference image to generated frames, improving transition smoothness. Human judges overwhelmingly favor FLASH, with 65.78\% of 488 responses prefer FLASH over baselines. We strongly recommend watching the videos in the website: https://lihaoxin05.github.io/human_action_animation/, as motion artifacts are hard to notice from images.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。