arXiv:2510.22199cs.CVcs.GR2025-10

构建3D场景中真实抓取动作数据集,提升人机交互真实性

MOGRAS: Human Motion with Grasping in 3D Scenes

  • 构建大规模3D室内场景中人体抓取动作数据集
  • 验证现有方法在场景感知上的不足,提出有效适配方案
  • 适合虚拟现实与机器人交互研究者使用

生成与物体互动的逼真人体全身运动对机器人、虚拟现实和人机交互至关重要。现有方法虽能生成3D场景中的全身动作,但缺乏精细抓取的保真度;而专注于精确抓取的方法又常忽略周围3D环境。为解决这一难题,我们提出MOGRAS(Human MOtion with GRAsping in 3D Scenes),一个大规模数据集,包含丰富标注的3D室内场景中预抓取的全身行走动作与最终抓取姿态。我们利用MOGRAS对现有全身体力抓取方法进行基准测试,揭示其在场景感知上的局限性,并提出一种简单有效的适配方法,使已有模型能无缝融入3D场景。通过大量定量与定性实验,验证了数据集有效性及所提方法显著提升,为更真实的真人-场景交互铺平道路。

原文摘要 · Abstract (English)

Generating realistic full-body motion interacting with objects is critical for applications in robotics, virtual reality, and human-computer interaction. While existing methods can generate full-body motion within 3D scenes, they often lack the fidelity for fine-grained tasks like object grasping. Conversely, methods that generate precise grasping motions typically ignore the surrounding 3D scene. This gap, generating full-body grasping motions that are physically plausible within a 3D scene, remains a significant challenge. To address this, we introduce MOGRAS (Human MOtion with GRAsping in 3D Scenes), a large-scale dataset that bridges this gap. MOGRAS provides pre-grasping full-body walking motions and final grasping poses within richly annotated 3D indoor scenes. We leverage MOGRAS to benchmark existing full-body grasping methods and demonstrate their limitations in scene-aware generation. Furthermore, we propose a simple yet effective method to adapt existing approaches to work seamlessly within 3D scenes. Through extensive quantitative and qualitative experiments, we validate the effectiveness of our dataset and highlight the significant improvements our proposed method achieves, paving the way for more realistic human-scene interactions.

动作生成3D场景抓取模拟数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。