让机器人在物品乱动时仍能靠语言理解自适应调整任务
LangPert: Detecting and Handling Task-level Perturbations for Robust Object Rearrangement
- 用视觉语言模型监控环境与动作执行,识别任务扰动
- 通过分层思维链让大模型生成修正后的操作计划
- 在未知扰动下仍能提升完成率和执行效率
物体重排任务可能受任务级扰动(TLP)影响,即物品意外增加、移除或移动,会破坏视觉策略并根本性干扰任务可行性与进展。为应对这一挑战,我们提出 LangPert——一种基于语言的框架,用于检测并缓解桌面上重排任务中的 TLP 情况。LangPert 利用视觉语言模型(VLM)全面监控策略执行与环境 TLP,同时借助分层思维链(HCoT)推理机制增强大语言模型(LLM)的上下文理解能力,生成自适应的修正操作计划。实验表明,相比基线方法,LangPert 在多种 TLP 场景下表现更优,实现更高任务完成率、更好执行效率,并具备对未见场景的泛化潜力。
原文摘要 · Abstract (English)
Task execution for object rearrangement could be challenged by Task-Level Perturbations (TLP), i.e., unexpected object additions, removals, and displacements that can disrupt underlying visual policies and fundamentally compromise task feasibility and progress. To address these challenges, we present LangPert, a language-based framework designed to detect and mitigate TLP situations in tabletop rearrangement tasks. LangPert integrates a Visual Language Model (VLM) to comprehensively monitor policy's skill execution and environmental TLP, while leveraging the Hierarchical Chain-of-Thought (HCoT) reasoning mechanism to enhance the Large Language Model (LLM)'s contextual understanding and generate adaptive, corrective skill-execution plans. Our experimental results demonstrate that LangPert handles diverse TLP situations more effectively than baseline methods, achieving higher task completion rates, improved execution efficiency, and potential generalization to unseen scenarios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。