arXiv:2503.05114cs.ROcs.AI2025-03被引 4

用状态机生成动作序列,提升机器人语言操作的成功率

Look Before You Leap: Using Serialized State Machine for Language Conditioned Robotic Manipulation

  • 用串联状态机自动生成动作演示,避免遗漏场景
  • 在复杂任务中成功率高达98%,远超原有方法的60%
  • 适合需要长序列精确操作的机器人任务研究

面向语言驱动的机器人操作,现有模仿学习框架依赖示范数据覆盖所有情况,若未涵盖则易引发级联错误。为此,我们提出基于串联有限状态机(Serialized FSM)的框架,用于生成示范并提升长序列精准交互任务的成功率。通过环境动态变化且时序较长的拼图任务进行验证,实验结果显示,本方法在相关任务中的成功率达98%,而对比组使用现有方法最高仅达60%,部分任务几乎无法完成。

原文摘要 · Abstract (English)

Imitation learning frameworks for robotic manipulation have drawn attention in the recent development of language model grounded robotics. However, the success of the frameworks largely depends on the coverage of the demonstration cases: When the demonstration set does not include examples of how to act in all possible situations, the action may fail and can result in cascading errors. To solve this problem, we propose a framework that uses serialized Finite State Machine (FSM) to generate demonstrations and improve the success rate in manipulation tasks requiring a long sequence of precise interactions. To validate its effectiveness, we use environmentally evolving and long-horizon puzzles that require long sequential actions. Experimental results show that our approach achieves a success rate of up to 98 in these tasks, compared to the controlled condition using existing approaches, which only had a success rate of up to 60, and, in some tasks, almost failed completely.

机器人操作状态机模仿学习语言控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。