arXiv:2604.10096cs.CV2026-04被引 5

让机器人能长期协作自进化,从指令到动作全程闭环控制。

ABot-Claw: A Foundation for Persistent, Cooperative, and Self-Evolving Robotic Agents

论文配图:ABot-Claw: A Foundation for Persistent, Cooperative, and Self-Evolving Robotic Agents
图 1 · 摘自论文原文
  • 用能力驱动调度统一异构机器人接口,实现多机协同。
  • 视觉主导跨实体记忆保留上下文,支持长时间任务追踪。
  • 通过通用奖励模型实时反馈纠错,适合开放环境自演进系统。

当前具身智能系统在开放世界中仍存在高层推理与底层物理执行之间的显著差距。尽管视觉-语言-动作(VLA)模型具备强大的感知与直觉响应能力,但其开环特性限制了长时程表现。引入系统2认知机制的代理虽提升规划能力,却通常局限于预设工具包的封闭沙箱,难以控制真实系统。OpenClaw提供具有完整系统权限的本地化运行时,但缺乏支撑长期、多机器人执行的具身控制架构。为此,我们提出ABot-Claw,作为OpenClaw的具身扩展,集成:1)以能力驱动调度的统一具身接口,实现异构机器人协调;2)以视觉为中心的跨具身多模态记忆,实现持久上下文保留与精准检索;3)基于评判器的闭环反馈机制,结合通用奖励模型实现在线进度评估、局部修正与重规划。采用分层解耦架构,涵盖OpenClaw层、共享服务层与机器人具身层,支持真实世界交互,打通自然语言意图到物理动作的闭环,赋能开放动态环境中持续自演进的机器人代理。

原文摘要 · Abstract (English)

Current embodied intelligent systems still face a substantial gap between high-level reasoning and low-level physical execution in open-world environments. Although Vision-Language-Action (VLA) models provide strong perception and intuitive responses, their open-loop nature limits long-horizon performance. Agents incorporating System 2 cognitive mechanisms improve planning, but usually operate in closed sandboxes with predefined toolkits and limited real-system control. OpenClaw provides a localized runtime with full system privileges, but lacks the embodied control architecture required for long-duration, multi-robot execution. We therefore propose ABot-Claw, an embodied extension of OpenClaw that integrates: 1) a unified embodiment interface with capability-driven scheduling for heterogeneous robot coordination; 2) a visual-centric cross-embodiment multimodal memory for persistent context retention and grounded retrieval; and 3) a critic-based closed-loop feedback mechanism with a generalist reward model for online progress evaluation, local correction, and replanning. With a decoupled architecture spanning the OpenClaw layer, shared service layer, and robot embodiment layer, ABot-Claw enables real-world interaction, closes the loop from natural language intent to physical action, and supports progressively self-evolving robotic agents in open, dynamic environments.

机器人具身智能自进化多机协作

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。