arXiv:2603.23676cs.AIcs.RO2026-03

用3D视觉语言模型实现长程物品重排,自动理解指令并规划动作。

Grounding Vision and Language to 3D Masks for Long-Horizon Box Rearrangement

  • 通过3D掩码预测实现语义与空间的联合推理,动态生成拾取放置动作。
  • 在含1-30个盒子的复杂场景中达成79.5%的成功率,显著优于2D模型方法。
  • 适合需要理解复杂自然语言指令的机器人长程任务场景。

我们研究仅使用视觉观测和不明确的自然语言目标,在3D环境中进行长程规划,聚焦多步3D箱子重排任务。现有方法通常依赖符号规划器,其状态与目标的关系绑定脆弱,或直接由2D视觉语言模型(VLM)生成动作序列,二者均难以处理多物体、丰富的3D几何结构及隐式语义约束。近期3D VLM在自然语言指代与3D分割掩码间的强对齐能力表明其具备更通用的规划潜力。我们扩展现有3D对齐模型,提出反应式动作掩码规划器(RAMP-3D),将长程规划建模为连续的3D掩码对预测:一个“选哪个物体”掩码与一个“放哪里”掩码。该系统基于RGB-D观测与自然语言任务描述,反应式生成3D箱体重排的多步拾取放置动作。我们在包含1-30个箱子的仓库式环境中测试了11种任务变体。RAMP-3D在长程重排任务中达到79.5%的成功率,显著优于基于2D VLM的基线,验证了基于掩码的反应式策略是长程规划中符号流水线的有力替代方案。

原文摘要 · Abstract (English)

We study long-horizon planning in 3D environments from under-specified natural-language goals using only visual observations, focusing on multi-step 3D box rearrangement tasks. Existing approaches typically rely on symbolic planners with brittle relational grounding of states and goals, or on direct action-sequence generation from 2D vision-language models (VLMs). Both approaches struggle with reasoning over many objects, rich 3D geometry, and implicit semantic constraints. Recent advances in 3D VLMs demonstrate strong grounding of natural-language referents to 3D segmentation masks, suggesting the potential for more general planning capabilities. We extend existing 3D grounding models and propose Reactive Action Mask Planner (RAMP-3D), which formulates long-horizon planning as sequential reactive prediction of paired 3D masks: a "which-object" mask indicating what to pick and a "which-target-region" mask specifying where to place it. The resulting system processes RGB-D observations and natural-language task specifications to reactively generate multi-step pick-and-place actions for 3D box rearrangement. We conduct experiments across 11 task variants in warehouse-style environments with 1-30 boxes and diverse natural-language constraints. RAMP-3D achieves 79.5% success rate on long-horizon rearrangement tasks and significantly outperforms 2D VLM-based baselines, establishing mask-based reactive policies as a promising alternative to symbolic pipelines for long-horizon planning.

3D规划视觉语言机器人掩码预测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。