arXiv:2410.02874cs.ROcs.AI2024-10中稿 · Advanced Robotics,…被引 11

用大模型和视觉语言模型实现真实厨房的自动烹饪

Real-World Cooking Robot System from Recipes Based on Food State Recognition Using Foundation Models and PDDL

  • 结合大模型与PDDL规划生成可执行烹饪动作
  • 仅用少量数据训练即可识别食材状态
  • 适用于需要灵活应对新菜谱的家用机器人

尽管机器人执行烹饪任务的需求日益增长,但在真实环境中基于新菜谱自主完成系列烹饪行为仍未能实现。本研究提出一种集成系统:利用大语言模型(LLM)生成可执行的烹饪行为规划,并通过经典规划语言PDDL描述动作;同时采用视觉-语言模型(VLM)在少量数据下学习食材状态识别。实验中,双臂轮式机器人PR2在真实厨房环境中成功执行了根据新菜谱安排的烹饪流程,验证了该系统的有效性。

原文摘要 · Abstract (English)

Although there is a growing demand for cooking behaviours as one of the expected tasks for robots, a series of cooking behaviours based on new recipe descriptions by robots in the real world has not yet been realised. In this study, we propose a robot system that integrates real-world executable robot cooking behaviour planning using the Large Language Model (LLM) and classical planning of PDDL descriptions, and food ingredient state recognition learning from a small number of data using the Vision-Language model (VLM). We succeeded in experiments in which PR2, a dual-armed wheeled robot, performed cooking from arranged new recipes in a real-world environment, and confirmed the effectiveness of the proposed system.

机器人烹饪大模型视觉语言模型PDDL

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。