arXiv:2607.26518cs.CV2026-07

构建首个第一人称视角安全理解评测基准,检验模型真推理能力

EgoSafe: A First-Person Mobile-Captured Benchmark for Visual Safety Understanding

论文配图:EgoSafe: A First-Person Mobile-Captured Benchmark for Visual Safety Understanding
图 1 · 摘自论文原文
  • 设计分层推理评估协议,强制模型完成从特征定位到意图推断的完整逻辑链
  • 在1.2万样本上测试顶级视觉语言模型,发现其描述准确但因果推理脆弱
  • 适合关注视频理解逻辑性、安全场景推理的研究者与工程师

真实场景中的可靠视觉安全理解不仅需要物体识别,更需在认知不确定性下进行因果推理。尽管大视觉语言模型(LVLMs)在标准基准上表现出色,但在第一人称动态、部分可观测的体验中,常无法区分表面相关与真实因果逻辑。现有评估多依赖第三人称监控视频和二分类指标,难以暴露这一认知差距。为此,我们提出EgoSafe-Bench,一个专为第一人称安全场景设计的评测基准,包含12,000个独特样本,由3,000段视频与按分层推理评估(HRE)协议生成的问答链配对构成。HRE要求模型从初始特征锚定出发,经历盲区推断与意图推理,确保逻辑一致性并惩罚捷径预测。对Qwen3-VL、Gemini、VideoLLaMA 3等主流模型的评估显示显著的感知-推理脱节:模型虽有高描述得分,但在因果推理与逻辑闭环上表现脆弱。本工作提供挑战性数据集与系统化评估框架,推动具备逻辑鲁棒性的视频理解系统发展。

原文摘要 · Abstract (English)

Reliable visual safety understanding in real-world scenarios demands more than just object recognition; it requires causal reasoning under epistemic uncertainty. While Large Vision-Language Models (LVLMs) demonstrate impressive semantic alignment on standard benchmarks, they often struggle to distinguish between superficial correlation and genuine forensic logic when grounded in the dynamic, partially observable nature of first-person experiences. Existing evaluations, dominated by third-person surveillance footage and binary classification metrics, fail to expose this cognitive gap. To address this, we introduce EgoSafe-Bench, a benchmark specifically designed to probe forensic reasoning in egocentric safety scenarios. It comprises 12,000 unique evaluation samples, generated by pairing each of the 3,000 video clips with a QA chain governed by our proposed Hierarchical Reasoning Evaluation (HRE) protocol. Unlike standard benchmarks, HRE mandates a rigorous reasoning trajectory from initial feature anchoring to blind-spot deduction and intent inference, thereby enforcing logical consistency and penalizing shortcut-based predictions. Extensive evaluations of state-of-the-art LVLMs (e.g., Qwen3-VL, Gemini, VideoLLaMA 3) reveal a significant perception-reasoning decoupling: models often achieve high descriptive scores but exhibit notable fragility in causal reasoning and logical closure. Our work provides both a challenging dataset and a systematic evaluation framework to foster the development of logically robust video understanding systems.

视觉推理第一人称视频评测基准因果推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。