让机器人在真实环境中长期自主执行复杂任务。
Advancing Omnimodal Embodied Agents from Isolated Skills to Everyday Physical Autonomy

- 分层异步架构分离规划、记忆与验证,实现跨虚拟与物理动作的统一控制。
- 实测40个长期任务中成功率全面提升,上下文增长接近线性,耗能稳定。
- 适合研究长期自主机器人系统或智能体框架的开发者参考。
在非结构化环境中构建持久存在的具身智能体,需要统一协调涵盖网络(API、IoT)与物理(操作、导航)领域的异构工具,并具备在长期运行中自主应对物理故障的能力。现有系统将这些问题割裂处理:基于视觉语言模型(VLM)的规划器缺乏统一的虚实动作空间,智能体框架累积无界上下文导致时间连贯性下降,而视觉语言行动(VLA)策略以开环方式执行且无法检测自身失败。我们主张持久自主应依赖分层异步架构,明确分离规划、记忆与验证。为此,提出OmniAct框架:整合多模态语义规划器用于跨统一动作空间的技能调度,采用事件边界驱动的自适应分层记忆实现次线性上下文增长,以及异步视觉预占引擎在物理执行中闭环验证语义。在两台机器人平台、四类物联网设备上完成40个真实世界长周期任务,OmniAct在所有复杂度层级均实现端到端成功率提升,累积交互超10万次时仍保持近似平坦的令牌消耗,并使中等规模开源模型性能达到专有水平。
原文摘要 · Abstract (English)
Building persistent embodied agents in unstructured environments demands unified orchestration of heterogeneous tools spanning both cyber (APIs, IoT) and physical (manipulation, navigation) domains, coupled with autonomous recovery from physical failures that inevitably arise over extended operation. Existing systems treat these as separate problems: VLM-based planners lack a unified cyber-physical action space, agent frameworks accumulate unbounded context that degrades temporal coherence, and VLA policies execute open-loop without detecting their own failures. We argue that persistent autonomy requires not a monolithic model but a hierarchical asynchronous architecture with explicit separation of planning, memory, and verification. To this end, we present OmniAct, a framework integrating a multimodal semantic planner for skill routing across unified action spaces, an adaptive hierarchical memory with event-boundary-driven compression for sub-linear context growth, and an asynchronous visual preemption engine that closes the semantic loop during physical execution. Across 40 real-world long-horizon tasks on two robotic platforms coordinating four IoT devices, OmniAct achieves consistent improvements in end-to-end success across all complexity levels, maintains near-flat token consumption over under 100k+ accumulated interaction tokens, and elevates mid-scale open-weight models to proprietary-level performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。