arXiv:2505.03035cs.ROcs.AI2025-05被引 13

用场景图提升语言模型在复杂环境中的机器人重排规划能力

MORE: Mobile Manipulation Rearrangement Through Grounded Language Reasoning

  • 通过场景图与实例区分,构建可约束的规划问题
  • 在BEHAVIOR-1K上解决81个任务中的大量难题,首次突破基准
  • 适用于室内外真实场景,支持日常活动类任务

自主长时程移动操作面临场景动态、未知区域和错误恢复等挑战。现有基于基础模型的方法在处理大量物体和大规模环境时性能下降。为此,我们提出MORE,一种增强语言模型在零样本条件下完成重排任务规划的新方法。MORE利用场景图表示环境,引入实例区分,并设计主动过滤机制,提取任务相关的对象与区域子图,使规划问题保持可控,有效减少幻觉并提升可靠性。此外,我们引入多项改进,使系统可跨室内外环境运行。我们在BEHAVIOR-1K基准的81个多样化重排任务上评估,MORE成为首个成功解决其中大量任务的方法,显著优于近期基于基础模型的方案。我们还在多个复杂真实任务中验证其能力,模拟日常行为。代码已公开于https://more-model.cs.uni-freiburg.de。

原文摘要 · Abstract (English)

Autonomous long-horizon mobile manipulation encompasses a multitude of challenges, including scene dynamics, unexplored areas, and error recovery. Recent works have leveraged foundation models for scene-level robotic reasoning and planning. However, the performance of these methods degrades when dealing with a large number of objects and large-scale environments. To address these limitations, we propose MORE, a novel approach for enhancing the capabilities of language models to solve zero-shot mobile manipulation planning for rearrangement tasks. MORE leverages scene graphs to represent environments, incorporates instance differentiation, and introduces an active filtering scheme that extracts task-relevant subgraphs of object and region instances. These steps yield a bounded planning problem, effectively mitigating hallucinations and improving reliability. Additionally, we introduce several enhancements that enable planning across both indoor and outdoor environments. We evaluate MORE on 81 diverse rearrangement tasks from the BEHAVIOR-1K benchmark, where it becomes the first approach to successfully solve a significant share of the benchmark, outperforming recent foundation model-based approaches. Furthermore, we demonstrate the capabilities of our approach in several complex real-world tasks, mimicking everyday activities. We make the code publicly available at https://more-model.cs.uni-freiburg.de.

机器人规划语言模型场景图重排任务

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。