arXiv:2412.13185cs.CV2024-12CVPR被引 10

用场景图生成适配环境的人类动作,让虚拟人物动得更自然。

Move-in-2D: 2D-Conditioned Human Motion Generation

  • 用扩散模型结合场景图和文本提示生成动作序列
  • 在真实场景图上生成的动作与场景投影对齐度高
  • 适合需要人景协调的视频生成任务

生成逼真人类视频仍是难题,现有方法多依赖已有动作序列作为控制信号,受限于特定动作类型和全局场景匹配。本文提出Move-in-2D,一种基于场景图像条件生成人类动作序列的新方法,使动作能自适应不同场景。该方法采用扩散模型,以场景图像和文本提示为输入,生成与场景契合的动作序列。为训练模型,我们构建了一个大规模单人活动视频数据集,并标注每个视频对应的人体运动作为目标输出。实验表明,所生成的动作在投影后与场景图像高度一致,且显著提升了视频合成中的人体动作质量。

原文摘要 · Abstract (English)

Generating realistic human videos remains a challenging task, with the most effective methods currently relying on a human motion sequence as a control signal. Existing approaches often use existing motion extracted from other videos, which restricts applications to specific motion types and global scene matching. We propose Move-in-2D, a novel approach to generate human motion sequences conditioned on a scene image, allowing for diverse motion that adapts to different scenes. Our approach utilizes a diffusion model that accepts both a scene image and text prompt as inputs, producing a motion sequence tailored to the scene. To train this model, we collect a large-scale video dataset featuring single-human activities, annotating each video with the corresponding human motion as the target output. Experiments demonstrate that our method effectively predicts human motion that aligns with the scene image after projection. Furthermore, we show that the generated motion sequence improves human motion quality in video synthesis tasks.

动作生成扩散模型场景适配

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。