用视觉语言模型自动从演示中生成机器人长程操作的符号规划描述
PDDL-ART: Autonomous Symbolic Abstraction From Demonstration For Long-Horizon Robotic Manipulation Using Vision-Language Models

- 基于视觉语言模型,仅凭一次示范和自然语言即可自动生成PDDL任务描述
- 在复杂家务与机械维护任务中达成93.3%成功率,优于基线15个百分点
- 通过几何与时间推理增强语义对齐,适合需抽象推理的长周期操作场景
符号规划中的PDDL为长程机器人操作提供了严谨框架,但构建准确的PDDL领域与问题描述仍存在显著瓶颈,通常需大量领域知识。我们提出一种基于视觉语言模型(VLM)的方法PDDL-ART,可从单次专家示范、自然语言任务描述及高阶动作名称库中自主生成特定任务的PDDL领域与问题描述。PDDL-ART无需任何领域模板、动作签名或微调。为确保生成内容不仅语法正确且语义契合示范任务,该框架引入多阶段纠错流程,涵盖语法、语义与执行层面。执行引导纠错的关键是符号谓词接地:不依赖单一视觉观察,而是利用现代VLM的工具使用能力,结合几何与时间推理来评估无法仅从图像判断的关系谓词。关键在于模型能自主决定何时调用工具并解释其输出。我们在发动机维护与家庭场景的复杂操作任务上评估PDDL-ART,包括需要记忆、抽象谓词推断以及目标状态与初始状态视觉不可区分的任务。PDDL-ART平均成功率达93.3%,高于基线VLM规划器的78.3%。
原文摘要 · Abstract (English)
Symbolic planning with PDDL offers a principled framework for long-horizon robot manipulation, but constructing accurate PDDL domain and problem descriptions remains a significant bottleneck, typically requiring substantial domain expertise. We present a Vision-Language Model (VLM)-based approach called PDDL-ART, a framework that autonomously generates task-specific PDDL domain and problem descriptions from a single expert demonstration, a natural language task description, and a library of available high-level action names. PDDL-ART does not require any domain templates, action signatures, or fine-tuning. To ensure the generated descriptions are not only syntactically valid but semantically aligned with the demonstrated task, PDDL-ART introduces a multi-stage correction pipeline operating at syntactic, semantic, and execution levels. A key component of execution-guided correction is symbolic predicate grounding. Instead of relying solely on visual observations, PDDL-ART leverages the tool-use capabilities of modern VLMs to incorporate geometric and temporal reasoning for evaluating relational predicates that are not directly discernible from images alone. Critically, the model autonomously determines when to invoke these tools and how to interpret their outputs. We evaluate PDDL-ART on challenging manipulation tasks in engine maintenance and household domains, including tasks that require memory, abstract predicate inference, and goal states that are visually indistinguishable from the initial state. PDDL-ART achieves an average success rate of 93.3%, compared to 78.3% for a baseline VLM-based planner.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。