arXiv:2501.17665cs.ROcs.AI2025-01被引 6

用视觉语言模型自动把图像转成机器人规划代码,降低使用门槛。

Planning with Vision-Language Models and a Use Case in Robot-Assisted Teaching

  • 通过视觉语言模型将图像和文字描述转为PDDL规划问题
  • 在积木世界等任务中生成语法与内容均正确的规划实例
  • 适合希望快速构建复杂规划任务的开发者或教育应用

利用大语言模型自动生成规划领域定义语言(PDDL)开启了人工智能规划的新方向,尤其适用于复杂现实任务。本文提出Image2PDDL框架,借助视觉语言模型(VLMs)将初始状态图像和目标状态描述自动转换为PDDL问题。通过提供配套的PDDL领域定义,该方法有效解决感知理解与符号规划之间的鸿沟,降低创建结构化问题实例的专家依赖,并提升跨不同复杂度任务的可扩展性。我们在多个领域(包括积木世界、滑动拼图等标准规划域)上评估该框架,使用多难度等级的数据集进行测试。评估指标涵盖语法正确性(确保语法规范且可执行)和内容正确性(验证生成的PDDL是否准确表示状态)。结果表明,该方法在多种任务复杂度下均表现良好,具备向更广泛应用拓展的潜力。我们还将探讨其在自闭症谱系障碍学生机器人辅助教学中的潜在应用场景。

原文摘要 · Abstract (English)

Automating the generation of Planning Domain Definition Language (PDDL) with Large Language Model (LLM) opens new research topic in AI planning, particularly for complex real-world tasks. This paper introduces Image2PDDL, a novel framework that leverages Vision-Language Models (VLMs) to automatically convert images of initial states and descriptions of goal states into PDDL problems. By providing a PDDL domain alongside visual inputs, Imasge2PDDL addresses key challenges in bridging perceptual understanding with symbolic planning, reducing the expertise required to create structured problem instances, and improving scalability across tasks of varying complexity. We evaluate the framework on various domains, including standard planning domains like blocksworld and sliding tile puzzles, using datasets with multiple difficulty levels. Performance is assessed on syntax correctness, ensuring grammar and executability, and content correctness, verifying accurate state representation in generated PDDL problems. The proposed approach demonstrates promising results across diverse task complexities, suggesting its potential for broader applications in AI planning. We will discuss a potential use case in robot-assisted teaching of students with Autism Spectrum Disorder.

AI规划视觉语言模型机器人教学

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。