arXiv:2606.23565cs.ROcs.CV2026-06被引 2

让机器人像人一样在真实世界中持续思考、记忆和执行复杂任务。

HoloAgent-0: A Unified Embodied Agent Framework with 3D Spatial Memory

论文配图:HoloAgent-0: A Unified Embodied Agent Framework with 3D Spatial Memory
图 1 · 摘自论文原文
  • 用三维空间记忆实现机器人对物理世界的长期感知与定位
  • 支持跨机器人协作与长时间任务的闭环执行,成功率超90%
  • 适合需要真实世界部署的智能机器人系统研发人员

大语言模型代理在数字环境中遵循推理-调用工具-反馈修正的循环。将其扩展到物理机器人面临连续执行、具身依赖、不确定性及安全约束等挑战。现有系统虽具备操作、空间理解、导航和拟人控制能力,但多为孤立模块或松散耦合。本文提出HoloAgent-0,一个面向真实机器人部署的统一具身智能体框架。Embodied AgentOS将语言指令转化为可执行技能图,调度资源、监控运行,并根据实时反馈触发澄清或重规划。该框架通过三层耦合结构整合异构机器人:具身代理操作系统(闭环执行)、三维空间记忆(物理世界锚定)与具身技能(动作执行)。我们在真实硬件上部署并评估了其空间记忆、长时程导航与闭环执行能力,涵盖运动生成、物体搜索、跨机器人协同与移动操作任务。

原文摘要 · Abstract (English)

LLM agents follow a practical execution loop in digital environments: they reason over structured states, invoke tools, inspect feedback, and revise actions. Extending this loop to physical robots is difficult because physical execution is continuous, embodiment-dependent, uncertain, and constrained by safety. Existing embodied-AI systems have advanced manipulation, spatial understanding, navigation, and humanoid control, but these capabilities often remain specialized modules or loosely coupled decision loops. In this work, we introduce HoloAgent-0, a unified embodied agent framework for real-world robot deployment. Embodied AgentOS converts language instructions into executable skill graphs, schedules robot resources, monitors execution, and triggers clarification or re-planning from runtime feedback. HoloAgent-0 organizes heterogeneous robot models and controllers through three coupled layers: Embodied AgentOS for closed-loop execution, 3D spatial memory for physical world grounding, and embodied skills for robot action. We deploy HoloAgent-0 on real hardware and evaluate its spatial memory, long-horizon navigation, and closed-loop execution across motion generation, object search, cross-robot coordination, and mobile manipulation.

具身智能机器人空间记忆任务规划

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。