arXiv:2602.00593cs.CVcs.LG2026-02被引 1

评测视觉语言模型在真实场景中细粒度理解与外部知识结合的能力

Pix2Fact: When Vision Is Not Enough -- Benchmarking Fine-Grained VQA with Web Verification on High-Resolution Real-World Scenes

  • 构建高分辨率图像问答数据集,要求结合视觉定位与外部知识
  • 顶级模型最高仅51.7%准确率,远低于人类专家水平
  • 适合研究多模态推理、知识融合与真实世界应用的学者

尽管通用任务取得进展,视觉语言模型在需要精细视觉定位与外部知识融合的挑战性任务上仍表现不佳,而现有基准未能同时评估这两项能力。为此,我们提出Pix2Fact,一个面向专家级视觉感知与知识检索的视觉问答基准。该数据集包含1000张4K以上高分辨率图像,覆盖八个真实场景,问题与答案由全球顶尖高校博士级标注者跨学科精心设计。每道题需精确视觉定位并整合外部知识。我们评估了十种最先进的视觉语言模型,包括Gemini-3.1-Pro和GPT-5.4等专有模型,发现其表现受限:即使提供视觉真值与搜索工具,最先进模型平均准确率仅为51.7%。分析表明,低准确率主要源于三方面:即便有视觉真值仍频繁出现定位错误、搜索利用浅层、无法获取长尾且非结构化的本地信息。这一显著差距揭示了当前模型在应对复杂现实场景时的局限性。我们认为Pix2Fact将推动下一代无缝融合精细感知与强知识检索的语言-视觉智能体的发展。

原文摘要 · Abstract (English)

Despite progress on general tasks, vision-language models (VLMs) still struggle with challenges that demand both fine-grained visual grounding and external knowledge, a synergy overlooked by existing benchmarks that evaluate these abilities in isolation. To fill this void, we introduce Pix2Fact, a visual question-answering benchmark designed to assess expert-level visual perception and knowledge search. Pix2Fact comprises 1,000 high-resolution (4K+) images spanning eight scenarios. Its questions and answers are meticulously crafted by PhD-holding annotators from top global universities across diverse disciplines. Each question requires detailed visual grounding and the integration of external knowledge. Evaluating ten state-of-the-art VLMs, including proprietary models such as Gemini-3.1-Pro and GPT-5.4, we find that Pix2Fact poses a formidable challenge: the most advanced model (Gemini-3.1-Pro) achieves only 51.7% average accuracy, even with access to visual ground truth and search tools. Our analysis attributes this low accuracy to three factors, frequent visual grounding errors even with visual ground truth, shallow search harnessing, and VLM's inability to retrieve long-tail, unstructured local information. This striking gap exposes the limitations of current models in assisting humans with real-world scenarios that demand overwhelming visual comprehension. We believe Pix2Fact will serve as a critical benchmark to drive the next generation of language-vision agents that seamlessly integrate fine-grained perception with robust knowledge search.

视觉问答多模态知识融合细粒度理解

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。