arXiv:2509.00767cs.CV2025-09中稿 · 3DV 2026被引 1

用网页视频自动构建大规模人物-物体交互数据集,提升动作生成真实感。

InterPose: Learning to Generate Human-Object Interactions from Large-Scale Web Videos

  • 从4.58万段网络视频中提取人与物体互动的3D动作
  • 构建含7.38万条动作序列的InterPose数据集,显著提升生成效果
  • 基于大模型实现零样本跨场景人形动画,适合影视游戏应用

人体动作生成近年来得益于大规模运动捕捉数据训练的扩散模型取得显著进展。然而,现有方法多聚焦于空场景中孤立人物的动画生成,而复杂三维场景中逼真人物-物体交互的合成仍是计算机图形学与机器人领域的关键挑战。主要障碍在于缺乏涵盖多样化物体操作的大规模数据集,现有运动捕捉数据通常仅限单人及有限物体的操作。为此,我们提出一套自动化动作提取流程,用于收集丰富的交互式人体动作数据。新构建的InterPose数据集包含从45.8万段视频中自动获取的73.8万条3D人体动作序列及其对应文本描述。通过大量实验验证,InterPose显著提升了当前先进的人体动作生成方法性能。此外,基于InterPose,我们开发了基于大语言模型的智能代理,实现了对多样化物体与场景的零样本人物动画生成。

原文摘要 · Abstract (English)

Human motion generation has shown great advances thanks to the recent diffusion models trained on large-scale motion capture data. Most of existing works, however, currently target animation of isolated people in empty scenes. Meanwhile, synthesizing realistic human-object interactions in complex 3D scenes remains a critical challenge in computer graphics and robotics. One obstacle towards generating versatile high-fidelity human-object interactions is the lack of large-scale datasets with diverse object manipulations. Indeed, existing motion capture data is typically restricted to single people and manipulations of limited sets of objects. To address this issue, we propose an automatic motion extraction pipeline and use it to collect interaction-rich human motions. Our new dataset InterPose contains 73.8K sequences of 3D human motions and corresponding text captions automatically obtained from 45.8K videos with human-object interactions. We perform extensive experiments and demonstrate InterPose to bring significant improvements to state-of-the-art methods for human motion generation. Moreover, using InterPose we develop an LLM-based agent enabling zero-shot animation of people interacting with diverse objects and scenes.

动作生成人物交互数据集扩散模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。