arXiv:2602.16682cs.CV2026-02被引 2

新基准SAW-Bench评估模型在真实场景中的自我感知能力

SAW-Bench: Learning Situated Awareness in the Real World

  • 用真实佩戴镜头视频构建视角中心的感知评测任务
  • 最佳模型Gemini 3 Flash仍比人类低37.66%表现
  • 适合研究视觉-动作协同与第一人称智能的学者

人类感知的核心是情境意识,即对自身所处物理环境的认知以及基于上下文进行行为推理的能力。然而,现有多模态基础模型(MFMs)评测大多聚焦场景内物体间的关系,忽视了需基于观察者视角、姿态和运动进行推理的主体中心关系。为此,我们提出SAW-Bench(Real World中的情境意识),一个基于真实世界视频的首人称情境意识评测基准。该基准包含786段由Ray-Ban Meta(Gen 2)智能眼镜自录的视频,覆盖多样室内外环境,以及超过2,071个由人工标注的问答对。它通过六类不同任务探测模型的主体中心理解能力。全面评估显示,即使使用表现最佳的MFM Gemini 3 Flash,仍存在37.66%的人机性能差距。深入分析发现:尽管模型可利用首人称视频中的部分几何线索,却常无法推断出一致的相机几何结构,导致系统性空间推理错误。SAW-Bench旨在推动对物理具身、以观察者为中心动态的理解,超越被动观测。

原文摘要 · Abstract (English)

A core aspect of human perception is situated awareness, the ability to relate ourselves to the surrounding physical environment and reason over possible actions in context. However, most existing benchmarks for multimodal foundation models (MFMs) emphasize environment-centric spatial relations (relations among objects in a scene), while largely overlooking observer-centric relationships that require reasoning relative to agent's viewpoint, pose, and motion. To bridge this gap, we introduce SAW-Bench (Situated Awareness in the Real World), a novel benchmark for evaluating egocentric situated awareness using real-world videos. SAW-Bench comprises 786 self-recorded videos captured with Ray-Ban Meta (Gen 2) smart glasses spanning diverse indoor and outdoor environments, and over 2,071 human-annotated question-answer pairs. It probes a model's observer-centric understanding with six different awareness tasks. Our comprehensive evaluation reveals a human-model performance gap of 37.66%, even with the best-performing MFM, Gemini 3 Flash. Beyond this gap, our in-depth analysis uncovers several notable findings; for example, while models can exploit partial geometric cues in egocentric videos, they often fail to infer a coherent camera geometry, leading to systematic spatial reasoning errors. We position SAW-Bench as a benchmark for situated spatial intelligence, moving beyond passive observation to understanding physically grounded, observer-centric dynamics.

情境意识第一人称视觉多模态评测空间推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。