统一视觉模型实现流式感知、重建与决策,支持实时交互。
OmniStream: Mastering Perception, Reconstruction and Action in Continuous Streams
- 用因果时空注意力+3D旋转位置编码,支持逐帧在线处理视频流。
- 在29个数据集上预训练,冻结主干仍媲美专用模型。
- 适合需要多模态理解的机器人和交互式智能体应用。
现代视觉智能体需具备通用、因果且物理结构化的表征能力,以应对实时流式环境。然而当前视觉基础模型仍分散,仅擅长图像语义感知、离线时序建模或空间几何重建。本文提出OmniStream,一种统一的流式视觉主干网络,能从多样化视觉输入中实现感知、重建与行动。通过引入因果时空注意力与3D旋转位置编码(3D-RoPE),模型借助持续的键值缓存(KV-cache)实现高效逐帧在线处理。我们在29个数据集上采用协同多任务框架进行预训练,联合优化静态与时序表征学习、流式几何重建及视觉-语言对齐。大量评估表明,即使主干完全冻结,OmniStream在图像与视频探测、流式几何重建、复杂视频与空间推理,以及机器人操控(未见于训练)任务中均表现优异,超越专用专家模型。本工作证明了训练单一通用视觉主干的可行性,该主干可跨语义、空间与时间推理泛化,是迈向通用视觉理解的重要一步。
原文摘要 · Abstract (English)
Modern visual agents require representations that are general, causal, and physically structured to operate in real-time streaming environments. However, current vision foundation models remain fragmented, specializing narrowly in image semantic perception, offline temporal modeling, or spatial geometry. This paper introduces OmniStream, a unified streaming visual backbone that effectively perceives, reconstructs, and acts from diverse visual inputs. By incorporating causal spatiotemporal attention and 3D rotary positional embeddings (3D-RoPE), our model supports efficient, frame-by-frame online processing of video streams via a persistent KV-cache. We pre-train OmniStream using a synergistic multi-task framework coupling static and temporal representation learning, streaming geometric reconstruction, and vision-language alignment on 29 datasets. Extensive evaluations show that, even with a strictly frozen backbone, OmniStream achieves consistently competitive performance with specialized experts across image and video probing, streaming geometric reconstruction, complex video and spatial reasoning, as well as robotic manipulation (unseen at training). Rather than pursuing benchmark-specific dominance, our work demonstrates the viability of training a single, versatile vision backbone that generalizes across semantic, spatial, and temporal reasoning, i.e., a more meaningful step toward general-purpose visual understanding for interactive and embodied agents.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。