用类人类流程规范提升机器人长序列操作的可靠性和自纠错能力
Agentic Robot: A Brain-Inspired Framework for Vision-Language-Action Models in Embodied Agents
- 采用标准化动作流程协调规划、执行与验证三阶段,模仿人类工作规程
- 在LIBERO基准上达成79.6%成功率,优于现有方法6.1%~7.4%的提升
- 适合需要高可靠性与可解释性的现实场景机器人系统开发
长时序机器人操作对自主系统提出严峻挑战,需具备持续推理、精准执行与鲁棒错误恢复能力。当前基于静态规划或端到端视觉运动策略的方法普遍存在误差累积问题,且执行中缺乏有效验证机制,限制其在真实场景中的可靠性。本文提出Agentic Robot——一种受大脑启发的框架,通过标准化动作流程(SAP)实现组件间的协同控制。SAP借鉴人类组织的标准化操作规程,构建了计划、执行与验证的结构化流程。系统包含三个专精模块:(1)大模型将高层指令分解为语义连贯的子目标;(2)视觉-语言-动作执行器从实时视觉输入生成连续控制命令;(3)时间验证器通过自我评估实现自主推进与错误恢复。该SAP驱动的闭环设计支持无需外部监督的动态自验证。在LIBERO基准上,Agentic Robot平均成功率达79.6%,较SpatialVLA提升6.1%,较OpenVLA提升7.4%,证明专用组件间基于SAP的协调显著提升了序列操作的性能与可解释性,展现出可靠自主系统的巨大潜力。
原文摘要 · Abstract (English)
Long-horizon robotic manipulation poses significant challenges for autonomous systems, requiring extended reasoning, precise execution, and robust error recovery across complex sequential tasks. Current approaches, whether based on static planning or end-to-end visuomotor policies, suffer from error accumulation and lack effective verification mechanisms during execution, limiting their reliability in real-world scenarios. We present Agentic Robot, a brain-inspired framework that addresses these limitations through Standardized Action Procedure (SAP)--a novel coordination protocol governing component interactions throughout manipulation tasks. Drawing inspiration from Standardized Operating Procedures (SOPs) in human organizations, SAP establishes structured workflows for planning, execution, and verification phases. Our architecture comprises three specialized components: (1) a large reasoning model that decomposes high-level instructions into semantically coherent subgoals, (2) a vision-language-action executor that generates continuous control commands from real-time visual inputs, and (3) a temporal verifier that enables autonomous progression and error recovery through introspective assessment. This SAP-driven closed-loop design supports dynamic self-verification without external supervision. On the LIBERO benchmark, Agentic Robot achieves state-of-the-art performance with an average success rate of 79.6%, outperforming SpatialVLA by 6.1% and OpenVLA by 7.4% on long-horizon tasks. These results demonstrate that SAP-driven coordination between specialized components enhances both performance and interpretability in sequential manipulation, suggesting significant potential for reliable autonomous systems. Project Github: https://agentic-robot.github.io.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。