构建首个面向多模态大模型流式空间智能的分级评测基准
OVO-S-Bench: A Hierarchical Benchmark for Streaming Spatial Intelligence in Multimodal LLMs

- 设计四层抽象层级,评估模型从感知到空间映射的连续推理能力
- 38个模型中最强者比人类差33分,全向空间映射是主要瓶颈
- 揭示思维链会放大无流证据支撑的空间错误,适合研究者参考
机器人、增强现实和自动驾驶中的多模态智能体需从持续的自我中心流中推断位置与布局,常依赖当前视野之外的信息。现有评测要么针对完整视频的离线评估,要么聚焦事件而非空间结构。本文提出OVO-S-Bench,一个完全人工标注的流式空间智能评测集,包含348个源视频上的1680个问题。标注由12名训练过的标注员完成(每人也担任盲评复核),历时约804人时的多轮质量保证。每个问题含查询时间戳与证据区间,评估时模型仅可见查询前缀。问题涵盖四个递增抽象层级:瞬时自我中心感知、时空上下文追踪、生成式空间推理、非中心空间映射。在38个专有及开源多模态大模型中,Gemini-3.1-Pro在相同前缀访问协议下比人类专家低33分(59.2 vs. 92.2),其中非中心空间映射为最大短板。值得注意的是,流式与空间微调模型表现反而低于其基础模型。此外发现,脱离流证据的思维链推理会加剧空间错误。该基准暴露了当前模型的关键缺陷,为下一代流式空间多模态大模型提供严苛测试平台。
原文摘要 · Abstract (English)
Multimodal agents in robotics, AR, and autonomous driving must reason about places and layouts from continuous egocentric streams, often using evidence outside the current view. Existing benchmarks either evaluate offline over full videos or target events rather than spatial structure. We introduce OVO-S-Bench, a fully human-annotated benchmark for streaming spatial intelligence, comprising 1,680 questions over 348 source videos. Annotation involves 12 trained annotators (each also serving as a blind cross-reviewer) across roughly 804 person-hours of multi-round quality assurance. Each question carries a query timestamp and an evidence interval, and at evaluation, the model sees only the prefix preceding the query. Questions span four levels of increasing abstraction: instantaneous egocentric perception, spatiotemporal context tracking, generative spatial reasoning, and allocentric spatial mapping. Across 38 proprietary and open-source MLLMs, Gemini-3.1-Pro trails human experts evaluated under the same prefix-access protocol by 33 points (59.2 vs. 92.2), with allocentric spatial mapping as the dominant bottleneck. Notably, streaming and spatially fine-tuned MLLMs underperform their own backbones. We further find that chain-of-thought reasoning amplifies spatial errors when ungrounded in the stream. By exposing these limitations, OVO-S-Bench establishes a demanding testbed for next-generation streaming spatial MLLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。