用领域相关的场景图提升机器人任务规划的准确性和成功率
Domain-Conditioned Scene Graphs for State-Grounded Task Planning
- 用领域条件化的场景图表示环境状态,结构更清晰
- 在三个领域中,状态识别准确率和任务成功率显著优于大模型方法
- 适合需要精确环境理解的机器人规划场景
近年来的机器人任务规划框架已整合如GPT-4o等大型多模态模型(LMM)。为解决这些模型的语义漂移问题,有研究建议将流程分为感知状态接地与后续的状态驱动规划。本文表明,基于LMM的状态接地能力仍受限于对细粒度、结构化、领域特定场景理解的不足。为此,我们提出一种更结构化的状态接地框架,采用领域条件化的场景图作为场景表示。该表示可直接映射至规划语言(如规划领域定义语言PDDL)中的符号状态,具备可操作性。我们实现了一个实例:通过轻量级视觉-语言方法,在领域相关的目标检测基础上分类领域特异性谓词,生成场景图。在三个领域的评估中,本方法在状态接地准确率和任务规划成功率上均显著优于基于LMM的方法。
原文摘要 · Abstract (English)
Recent robotic task planning frameworks have integrated large multimodal models (LMMs) such as GPT-4o. To address grounding issues of such models, it has been suggested to split the pipeline into perceptional state grounding and subsequent state-based planning. As we show in this work, the state grounding ability of LMM-based approaches is still limited by weaknesses in granular, structured, domain-specific scene understanding. To address this shortcoming, we develop a more structured state grounding framework that features a domain-conditioned scene graph as its scene representation. We show that such representation is actionable in nature as it is directly mappable to a symbolic state in planning languages such as the Planning Domain Definition Language (PDDL). We provide an instantiation of our state grounding framework where the domain-conditioned scene graph generation is implemented with a lightweight vision-language approach that classifies domain-specific predicates on top of domain-relevant object detections. Evaluated across three domains, our approach achieves significantly higher state rounding accuracy and task planning success rates compared to LMM-based approaches.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。