提出空间超感知新范式,让模型能预测、组织视觉经验。
Cambrian-S: Towards Spatial Supersensing in Video

- 构建两阶段评测框架VSI-SUPER,考验长期视频理解与空间推理。
- 在VSI-Bench上提升30%性能,但现有模型仍难突破空间建模瓶颈。
- 用预测误差驱动记忆分割,为下一代智能系统提供新思路。
我们认为,真正多模态智能的进步需要从反应式、任务驱动的系统转向更广泛的超感知范式。空间超感知包含四个阶段:语义感知(识别所见内容)、流式事件认知(维持连续体验的记忆)、隐式三维空间认知(推断像素背后的三维世界)和预测性世界建模(构建内部模型以过滤与组织信息)。当前基准大多仅测试早期阶段,对空间认知覆盖有限,且极少要求模型具备真正的世界建模能力。为此,我们提出VSI-SUPER,一个双任务基准:长时程视觉空间回忆(VSR)与持续视觉空间计数(VSC),支持任意长度视频输入,且不易被暴力扩展上下文所破解。通过构建VSI-590K数据集并训练Cambrian-S模型,在不牺牲通用能力的前提下,于VSI-Bench上实现+30%绝对性能提升。然而,其在VSI-SUPER上的表现仍受限,表明规模本身不足以达成空间超感知。我们提出预测感知作为前进路径,在自监督下一帧潜变量预测基础上,利用预测误差(惊喜度)驱动记忆与事件分割。在VSI-SUPER上,该方法显著优于领先专有基线,证明空间超感知需模型不仅‘看见’,还需‘预判’、‘选择’与‘组织’感知经验。
原文摘要 · Abstract (English)
We argue that progress in true multimodal intelligence calls for a shift from reactive, task-driven systems and brute-force long context towards a broader paradigm of supersensing. We frame spatial supersensing as four stages beyond linguistic-only understanding: semantic perception (naming what is seen), streaming event cognition (maintaining memory across continuous experiences), implicit 3D spatial cognition (inferring the world behind pixels), and predictive world modeling (creating internal models that filter and organize information). Current benchmarks largely test only the early stages, offering narrow coverage of spatial cognition and rarely challenging models in ways that require true world modeling. To drive progress in spatial supersensing, we present VSI-SUPER, a two-part benchmark: VSR (long-horizon visual spatial recall) and VSC (continual visual spatial counting). These tasks require arbitrarily long video inputs yet are resistant to brute-force context expansion. We then test data scaling limits by curating VSI-590K and training Cambrian-S, achieving +30% absolute improvement on VSI-Bench without sacrificing general capabilities. Yet performance on VSI-SUPER remains limited, indicating that scale alone is insufficient for spatial supersensing. We propose predictive sensing as a path forward, presenting a proof-of-concept in which a self-supervised next-latent-frame predictor leverages surprise (prediction error) to drive memory and event segmentation. On VSI-SUPER, this approach substantially outperforms leading proprietary baselines, showing that spatial supersensing requires models that not only see but also anticipate, select, and organize experience.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。