arXiv:2507.13405cs.CVcs.LG2025-07被引 4

构建首个聚焦密集场景推理的视觉问答基准,揭示现有模型在语义推理上的明显短板。

COREVQA: A Crowd Observation and Reasoning Entailment Visual Question Answering Benchmark

  • 基于CrowdHuman数据集生成5608对真假陈述,测试模型视觉蕴含推理能力
  • 顶尖模型准确率不足80%,多数模型仅39.98%-69.95%
  • 适合研究视觉推理、密集场景理解的学者与开发者参考

近年来,众多基准和数据集被用于通过视觉问答(VQA)对评估视觉语言模型(VLMs),模型性能显著提升。然而,这些基准很少测试模型在视觉蕴含任务中的表现,例如根据图像判断一个假设是否成立。为此,我们提出COREVQA(Crowd Observations and Reasoning Entailment),一个包含5608个图像与合成真假陈述对的基准,图像来自CrowdHuman数据集,旨在挑战模型在复杂拥挤场景下的视觉蕴含推理能力。实验结果表明,即使是最先进的VLMs,准确率也低于80%,其他模型表现更差(39.98%-69.95%)。这一显著差距揭示了当前模型在特定类型图像-问题对上推理能力的关键局限。

原文摘要 · Abstract (English)

Recently, many benchmarks and datasets have been developed to evaluate Vision-Language Models (VLMs) using visual question answering (VQA) pairs, and models have shown significant accuracy improvements. However, these benchmarks rarely test the model's ability to accurately complete visual entailment, for instance, accepting or refuting a hypothesis based on the image. To address this, we propose COREVQA (Crowd Observations and Reasoning Entailment), a benchmark of 5608 image and synthetically generated true/false statement pairs, with images derived from the CrowdHuman dataset, to provoke visual entailment reasoning on challenging crowded images. Our results show that even the top-performing VLMs achieve accuracy below 80%, with other models performing substantially worse (39.98%-69.95%). This significant performance gap reveals key limitations in VLMs' ability to reason over certain types of image-question pairs in crowded scenes.

视觉问答视觉推理密集场景语义蕴含

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。