arXiv:2603.22120cs.CV2026-03被引 2

StreamingClaw让智能体实现视频流实时感知决策闭环,支持长期记忆与主动交互。

StreamingClaw Technical Report

  • 构建支持实时流式推理与多模态长时记忆的统一框架
  • 实现感知-决策-行动闭环,支持未来事件预测与主动交互
  • 兼容OpenClaw生态,适合复杂物理环境下的智能体应用

新兴应用如具身智能、AI硬件、自动驾驶和智能座舱依赖于实时感知-决策-动作闭环,对视频流理解提出严苛挑战。当前智能体普遍存在能力碎片化问题,如仅支持离线视频理解、缺乏长期多模态记忆机制,或在流式输入下难以实现实时推理与主动交互。这些缺陷成为智能体在复杂现实环境中持续感知、实时决策与执行闭环动作的关键瓶颈,限制其在动态开放物理世界中的部署与潜力。为解决上述问题,我们提出StreamingClaw——一个面向视频流理解与具身智能的统一智能体框架。除保持与OpenClaw框架完全兼容外,它原生支持实时、多模态流式交互。StreamingClaw集成五大核心能力:(1) 支持实时流式推理;(2) 在交互目标在线演化下支持未来事件推理与主动交互;(3) 支持多模态长期记忆存储、层级记忆演化、高效记忆检索及多智能体间记忆共享;(4) 实现感知-决策-行动闭环,除传统工具与技能外,还提供适配真实物理环境的流式工具与以行为为中心的技能;(5) 兼容OpenClaw框架,可利用开源社区资源与支持。

原文摘要 · Abstract (English)

Emerging applications such as embodied intelligence, AI hardware, autonomous driving, and intelligent cockpits rely on a real-time perception-decision-action closed loop, posing stringent challenges for streaming video understanding. However, current agents mostly suffer from fragmented capabilities, such as supporting only offline video understanding, lacking long-term multimodal memory mechanisms, or struggling to achieve real-time reasoning and proactive interaction under streaming input. These shortcomings have become a key bottleneck for preventing agents from sustaining perception, making real-time decisions, and executing closed-loop actions in complex real-world environments, constraining their deployment and potential in dynamic, open physical worlds. To alleviate these issues, we propose StreamingClaw, a unified agent framework for streaming video understanding and embodied intelligence. Beyond maintaining full compatibility with the OpenClaw framework, it natively supports real-time, multimodal streaming interactions. StreamingClaw integrates five core capabilities: (1) It supports real-time streaming reasoning. (2) It supports reasoning about future events and proactive interaction under the online evolution of interaction objectives. (3) It supports multimodal long-term memory storage, hierarchical memory evolution, efficient memory retrieval, and memory sharing across multiple agents. (4) It supports a closed loop of perception-decision-action. In addition to conventional tools and skills, it also provides streaming tools and action-centric skills tailored for real-world physical environments. (5) It is compatible with the OpenClaw framework, allowing it to leverage the resources and support of the open-source community.

具身智能视频流智能体实时推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。