构建可长期运行的机器人操作系统,支持多模态记忆与自我进化。
ABot-AgentOS: A General Robotic Agent OS with Lifelong Multi-modal Memory

- 设计通用机器人操作系统,整合规划、记忆、验证与跨设备执行能力。
- 在16个场景200+任务上实现87.5%的LoCoMo准确率,支持持续改进。
- 通过故障驱动自进化机制防止数据泄露,适合复杂长期任务研究。
现有视觉语言模型与视觉语言动作系统虽提升了机器人感知与动作预测能力,但长时程具身智能体仍需通用运行时层以实现推理、记忆、工具使用、验证及跨平台执行。本文提出ABot-AgentOS,一个位于底层控制器之上的通用机器人代理操作系统,提供场景感知规划、上下文隔离技能执行、多阶段验证、多模态记忆与边云协同能力。为评估该系统,我们构建EmbodiedWorldBench,一个包含16个室内外混合场景、4种难度等级、超过200项任务的可执行基准,涵盖导航、物体搜寻、NPC对话、动态事件及基于轨迹的评分。ABot-AgentOS引入通用多模态图记忆,将对话、视觉观测、空间上下文、时间关系与任务轨迹转化为带类型节点与边的持久化结构。故障驱动的自进化循环将诊断出的记忆失败转化为有门控的运行时进化资产,仅在后续评估集推广,避免当前集真值泄漏的同时实现持续优化。在初始EmbodiedWorldBench子集上,ABot-AgentOS在任务成功率与目标完成度上均优于单控制器基线。在多个记忆基准上,静态版本达到LoCoMo 87.5、OpenEQA EM-EQA 59.9、Mem-Gallery 88.6、NExT-QA Acc@All 76.5;自进化进一步提升至LoCoMo 88.7、OpenEQA 60.4、Mem-Gallery 89.0。结果表明,通用代理操作系统能有效提升长时程具身执行能力,并提供可审计的持久记忆。
原文摘要 · Abstract (English)
Recent VLM and VLA systems have improved robotic perception and action prediction, yet long-horizon embodied agents still require a general runtime layer for reasoning, memory, tool use, verification, and cross-embodiment execution. We present ABot-AgentOS, a general robotic Agent Operating System that sits above low-level controllers and provides a deliberative agent layer for scene-conditioned planning, context-isolated skill execution, multi-stage verification, multi-modal memory, and edge-cloud collaboration. To evaluate such systems, we introduce EmbodiedWorldBench, an executable benchmark with 16 indoor, outdoor, and hybrid scenes, four difficulty levels, and over 200 tasks involving navigation, object search, NPC dialogue, dynamic events, and trace-grounded scoring. ABot-AgentOS further introduces Universal Multi-modal Graph Memory, a persistent source-grounded substrate that converts dialogue, visual observations, spatial context, temporal relations, and task traces into typed nodes and edges. A failure-driven self-evolution loop converts diagnosed memory failures into gated runtime evo-assets that are promoted only to later evaluation splits, preventing current-split ground-truth leakage while enabling continual improvement. On an initial EmbodiedWorldBench subset, ABot-AgentOS improves over a single-controller baseline in both task success and goal completion. Across memory benchmarks, ABot-AgentOS Static achieves 87.5 on LoCoMo, 59.9 on OpenEQA EM-EQA, 88.6 on Mem-Gallery, and 76.5 Acc@All on NExT-QA; self-evolution further improves LoCoMo to 88.7, OpenEQA to 60.4, and Mem-Gallery to 89.0. These results suggest that a general Agent OS layer can improve long-horizon embodied execution while providing persistent, auditable memory for continual interaction.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。