arXiv:2607.13681cs.CV2026-07中稿 · ECCV被引 1

构建真实场景长时视频基准,揭示模型空间推理能力的深层缺陷。

Towards Spatial Supersensing in the Wild

论文配图:Towards Spatial Supersensing in the Wild
图 1 · 摘自论文原文
  • 基于真实世界视频构建多模态长时空间感知评测集
  • 6980个问答对验证模型在4小时以上视频中追踪世界状态的能力
  • 发现模型存在空间坍塌、语义捷径等四类根本性失败模式

人类能高效解析持续数小时甚至数年的感官流,构建内化世界模型以支持空间推理与预测。为模拟此能力,空间超感知挑战多模态模型超越语言理解,实现真正世界建模。然而,现有基准依赖合成长视频(随机拼接短片段),且仅限家庭场景,缺乏真实世界的连续性与多样性。为此,我们提出VSI-Super-Wild,一个大规模基准,用于评估多样真实场景中长时程的空间超感知能力。受认知研究启发,系统探测世界状态三要素:主体(观察者)、物体(场景项)和环境(地点与全局布局)。该数据集包含6,980个经人工验证的问答对,源自442段真实视频,覆盖8类场景,部分视频长达4小时以上。实验结果揭示根本性断层:尽管静态图像理解取得进展,模型在需要长时间连贯世界状态追踪的任务中持续表现不佳。我们量化了性能随世界状态复杂度与时间跨度的退化,并诊断出四种失败模式:空间坍塌、语义捷径、更新不足与实例混淆。该分类表明,模型缺乏将物体、主体与环境整合为统一空间世界模型的机制,这是空间超感知发展的关键瓶颈。

原文摘要 · Abstract (English)

Humans can efficiently parse continuous sensory streams, from hours to years, scaffolding an internal world model that grounds spatial reasoning and prediction. To mimic this capacity, spatial supersensing challenges multimodal models to move beyond linguistic understanding toward true world modeling. However, their benchmark relies on synthetic long videos, formed by concatenating random short clips, and is mostly limited to household scenes, leaving real-world continuity and diversity underexplored. To address the gap, we introduce $\textbf{VSI-Super-Wild}$, a large-scale benchmark for evaluating spatial supersensing over long temporal horizons in diverse in-the-wild scenes. Notably, inspired by cognitive studies on how humans structure experience, we systematically probe the full triad of world state: the agent (observer), objects (scene items), and the environment (places and global layout). In total, VSI-Super-Wild contains $\textbf{6,980}$ human-verified question-answer pairs derived from $\textbf{442}$ real-world videos spanning 8 scene categories, including long-form recordings exceeding 4 hours. Results on VSI-Super-Wild expose a fundamental disconnect: despite advances in static image understanding, models consistently fail at tasks that require coherent world-state tracking over time. We characterize how performance degrades with world-state complexity and temporal horizon, and diagnose four failure modes: spatial collapse, semantic shortcuts, insufficient update, and instance confusion. This taxonomy reveals that models lack mechanisms to bind objects, agents, and environments into a unified spatial world model, a fundamental gap that defines the path forward for spatial supersensing.

空间建模长时序视觉理解多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。