arXiv:2504.03748cs.LGcs.AI2025-04

构建顶视图图像理解基准,评估视觉语言模型可靠性。

TDBench: A Benchmark for Top-Down Image Understanding with Reliability Analysis of Vision-Language Models

  • 设计旋转一致性评测方法,检验模型在不同视角下的回答稳定性。
  • 基于2000个问题×4种旋转,发现多数模型在旋转下表现不一致。
  • 适合关注AI可信性、导航与遥感场景的开发者和研究者。

顶视图像在自动驾驶和航拍等安全关键场景中至关重要,因其提供前视图无法捕捉的整体空间信息。然而,当前视觉语言模型(VLMs)主要在前视图基准上训练和评估,对其在顶视图下的性能了解不足。现有评估忽略了顶视图的一个关键特性:其物理意义在旋转下保持不变。此外,传统准确率指标易受幻觉或‘侥幸猜测’误导,掩盖了模型的真实可靠性与视觉证据基础。为此,我们提出TDBench,一个包含每场景2000个精心设计问题的顶视图理解基准,覆盖四种旋转角度。我们引入旋转一致性评测(RotationalEval, RE),衡量模型在相同场景不同旋转下的答案一致性,并建立可靠性框架,分离真实知识与偶然猜测。通过四个案例研究,揭示了模型在真实世界挑战中的薄弱环节。TDBench不仅为顶视图感知提供严谨评估,还从可信度角度推动更稳健、具象化的AI系统发展。

原文摘要 · Abstract (English)

Top-down images play an important role in safety-critical settings such as autonomous navigation and aerial surveillance, where they provide holistic spatial information that front-view images cannot capture. Despite this, Vision Language Models (VLMs) are mostly trained and evaluated on front-view benchmarks, leaving their performance in the top-down setting poorly understood. Existing evaluations also overlook a unique property of top-down images: their physical meaning is preserved under rotation. In addition, conventional accuracy metrics can be misleading, since they are often inflated by hallucinations or "lucky guesses", which obscures a model's true reliability and its grounding in visual evidence. To address these issues, we introduce TDBench, a benchmark for top-down image understanding that includes 2000 curated questions for each rotation. We further propose RotationalEval (RE), which measures whether models provide consistent answers across four rotated views of the same scene, and we develop a reliability framework that separates genuine knowledge from chance. Finally, we conduct four case studies targeting underexplored real-world challenges. By combining rigorous evaluation with reliability metrics, TDBench not only benchmarks VLMs in top-down perception but also provides a new perspective on trustworthiness, guiding the development of more robust and grounded AI systems. Project homepage: https://github.com/Columbia-ICSL/TDBench

视觉语言模型顶视图可靠性评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。