arXiv:2503.11081cs.ROcs.AI2025-03ICCV被引 16

构建10万+样本数据集,让机器人学会精准定位抓取物品

MoMa-Kitchen: A 100K+ Benchmark for Affordance-Grounded Last-Mile Navigation in Mobile Manipulation

  • 基于物体可操作性生成最优抓取位置标签
  • 支持不同机械臂和底盘高度的泛化定位能力
  • 适合研究具身智能与人机交互的开发者

在移动操作任务中,导航与操作常被割裂处理,导致仅靠近目标不足以有效执行操作。现有导航方法多以接近目标为成功标准,忽视了后续操作所需的最优姿态。为此,我们提出MoMa-Kitchen,一个包含超过10万样本的基准数据集,用于训练模型学习面向操作的最终导航位置。数据来自多样化的厨房环境,涵盖不同型号的移动操作机器人在杂乱场景中抓取目标物体的过程。通过全自动模拟流程,生成针对最优操作位置的可操作性标签。视觉数据由安装在机械臂上的第一视角RGB-D相机采集,保证视角一致性。我们还开发了轻量级基线模型NavAff,其在MoMa-Kitchen上表现优异。该方法使模型能够适应不同机械臂类型和平台高度,推动导航与操作更鲁棒、更通用的融合,助力具身人工智能发展。

原文摘要 · Abstract (English)

In mobile manipulation, navigation and manipulation are often treated as separate problems, resulting in a significant gap between merely approaching an object and engaging with it effectively. Many navigation approaches primarily define success by proximity to the target, often overlooking the necessity for optimal positioning that facilitates subsequent manipulation. To address this, we introduce MoMa-Kitchen, a benchmark dataset comprising over 100k samples that provide training data for models to learn optimal final navigation positions for seamless transition to manipulation. Our dataset includes affordance-grounded floor labels collected from diverse kitchen environments, in which robotic mobile manipulators of different models attempt to grasp target objects amidst clutter. Using a fully automated pipeline, we simulate diverse real-world scenarios and generate affordance labels for optimal manipulation positions. Visual data are collected from RGB-D inputs captured by a first-person view camera mounted on the robotic arm, ensuring consistency in viewpoint during data collection. We also develop a lightweight baseline model, NavAff, for navigation affordance grounding that demonstrates promising performance on the MoMa-Kitchen benchmark. Our approach enables models to learn affordance-based final positioning that accommodates different arm types and platform heights, thereby paving the way for more robust and generalizable integration of navigation and manipulation in embodied AI. Project page: \href{https://momakitchen.github.io/}{https://momakitchen.github.io/}.

移动操作可操作性具身智能数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。