通过文本生成可控人体动作,再合成高质量运动视频。
HumanDreamer: Generating Controllable Human-Motion Videos via Decoupled Generation
- 先用文本生成多样姿态,再基于姿态合成视频,解耦生成流程。
- 在MotionVid数据集上训练,文本到姿态生成准确率提升41.8%以上。
- 适合需要精准动作控制的视频生成、动作预测等任务使用。
人体动作视频生成面临核心挑战在于人体运动模式难以建模。现有方法虽尝试通过姿态控制驱动生成,但依赖已有视频提取姿态,灵活性不足。为此,我们提出HumanDreamer,一种解耦式人体视频生成框架:先从文本提示生成多样化姿态,再利用这些姿态生成人体运动视频。我们构建了最大规模的人体动作姿态生成数据集MotionVid;在此基础上,提出MotionDiT模型,可由文本生成结构化人体动作姿态,并引入新型LAMA损失,使FID指标下降62.4%,同时在top1、top2、top3的R-precision分别提升41.8%、26.3%、18.3%,显著提升文本到姿态的控制精度与生成质量。实验表明,本方法生成的姿态能驱动多种姿态到视频基线,生成多样且高质量的运动视频,且可拓展用于姿态序列预测和2D-3D动作还原等下游任务。
原文摘要 · Abstract (English)
Human-motion video generation has been a challenging task, primarily due to the difficulty inherent in learning human body movements. While some approaches have attempted to drive human-centric video generation explicitly through pose control, these methods typically rely on poses derived from existing videos, thereby lacking flexibility. To address this, we propose HumanDreamer, a decoupled human video generation framework that first generates diverse poses from text prompts and then leverages these poses to generate human-motion videos. Specifically, we propose MotionVid, the largest dataset for human-motion pose generation. Based on the dataset, we present MotionDiT, which is trained to generate structured human-motion poses from text prompts. Besides, a novel LAMA loss is introduced, which together contribute to a significant improvement in FID by 62.4%, along with respective enhancements in R-precision for top1, top2, and top3 by 41.8%, 26.3%, and 18.3%, thereby advancing both the Text-to-Pose control accuracy and FID metrics. Our experiments across various Pose-to-Video baselines demonstrate that the poses generated by our method can produce diverse and high-quality human-motion videos. Furthermore, our model can facilitate other downstream tasks, such as pose sequence prediction and 2D-3D motion lifting.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。