arXiv:2506.14428cs.CV2025-06被引 2

构建15万条双人互动动作视频数据集,实现高真实感2D人体动作生成

Toward Rich Video Human-Motion2D Generation

  • 基于扩散模型,融合全局与局部文本特征增强动作控制性
  • 双阶段训练:先标准扩散,再用FID奖励优化真实感与文本对齐
  • 首次支持复杂双人交互的2D动作生成,适合动画与虚拟角色开发

生成真实且可控的人体动作,尤其是涉及多人复杂互动的动作,仍因数据稀缺和人际动态建模复杂而面临挑战。为此,我们首先引入一个大规模的丰富视频人体动作2D数据集(Motion2D-Video-150K),包含15万条视频序列,涵盖均衡分布的单人动作与关键的双人交互动作,并配有详细文本描述。基于该数据集,我们提出一种新的基于扩散的丰富视频人体动作2D生成模型(RVHM2D)。RVHM2D采用改进的文本条件机制,使用双文本编码器(CLIP-L/B)或T5-XXL,结合全局与局部特征。我们设计了两阶段训练策略:先以标准扩散目标训练,再通过基于FID的强化学习进行微调,以进一步提升动作真实感与文本一致性。大量实验表明,RVHM2D在Motion2D-Video-150K基准上,于单人及双人交互场景下均达到领先性能。

原文摘要 · Abstract (English)

Generating realistic and controllable human motions, particularly those involving rich multi-character interactions, remains a significant challenge due to data scarcity and the complexities of modeling inter-personal dynamics. To address these limitations, we first introduce a new large-scale rich video human motion 2D dataset (Motion2D-Video-150K) comprising 150,000 video sequences. Motion2D-Video-150K features a balanced distribution of diverse single-character and, crucially, double-character interactive actions, each paired with detailed textual descriptions. Building upon this dataset, we propose a novel diffusion-based rich video human motion2D generation (RVHM2D) model. RVHM2D incorporates an enhanced textual conditioning mechanism utilizing either dual text encoders (CLIP-L/B) or T5-XXL with both global and local features. We devise a two-stage training strategy: the model is first trained with a standard diffusion objective, and then fine-tuned using reinforcement learning with an FID-based reward to further enhance motion realism and text alignment. Extensive experiments demonstrate that RVHM2D achieves leading performance on the Motion2D-Video-150K benchmark in generating both single and interactive double-character scenarios.

动作生成扩散模型双人互动2D动画

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。