arXiv:2503.21564cs.RO2025-03被引 2

用大模型+图网络生成可执行的烹饪任务计划,成功率提升至80%。

Cooking Task Planning using LLM and Verified by Graph Network

  • 结合大模型与图网络验证计划逻辑,减少幻觉错误。
  • 在5个食谱上测试,本方法成功执行4个计划,单一模型仅1个。
  • 适合需要高可靠性的机器人任务规划场景。

烹饪任务对机器人仍具挑战性,因其复杂性高。人类烹饪视频虽是宝贵信息源,但其表达多样,难以直接转化为机器人可执行指令。本文提出一种基于大语言模型(LLM)的任务与运动规划(TAMP)框架,从带字幕的烹饪视频中自动生成可执行的烹饪任务计划,并在双臂机器人上执行。传统基于LLM的方法因视频不确定性及输出幻觉而效果不佳。为此,本文引入功能对象导向网络(FOON)对计划进行验证并提供失败反馈,确保生成的动作序列逻辑正确且可执行。在5个食谱上的对比实验表明,本方法成功执行4个计划,而仅使用LLM的方法成功1个。

原文摘要 · Abstract (English)

Cooking tasks remain a challenging problem for robotics due to their complexity. Videos of people cooking are a valuable source of information for such task, but introduces a lot of variability in terms of how to translate this data to a robotic environment. This research aims to streamline this process, focusing on the task plan generation step, by using a Large Language Model (LLM)-based Task and Motion Planning (TAMP) framework to autonomously generate cooking task plans from videos with subtitles, and execute them. Conventional LLM-based task planning methods are not well-suited for interpreting the cooking video data due to uncertainty in the videos, and the risk of hallucination in its output. To address both of these problems, we explore using LLMs in combination with Functional Object-Oriented Networks (FOON), to validate the plan and provide feedback in case of failure. This combination can generate task sequences with manipulation motions that are logically correct and executable by a robot. We compare the execution of the generated plans for 5 cooking recipes from our approach against the plans generated by a few-shot LLM-only approach for a dual-arm robot setup. It could successfully execute 4 of the plans generated by our approach, whereas only 1 of the plans generated by solely using the LLM could be executed.

任务规划大模型机器人图网络

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。