挑战主流模型视觉理解能力,测试生成图像的细粒度多模态推理。
JourneyBench: A Challenging One-Stop Vision-Language Understanding Benchmark of Generated Images
- 构建生成图像的多任务基准,需细粒度跨模态推理。
- 顶尖模型在五项任务中表现均不理想,暴露出视觉推理短板。
- 适合评估模型在虚构场景下的真实理解能力,非依赖语言偏见。
现有视觉-语言理解基准多由物体在常规场景中的图像构成,导致当前多模态大模型仅通过浅层视觉理解即可借助语言背景偏见取得高分,因而性能与真实视觉理解能力无关。本文发布JourneyBench,一个全面的人工标注生成图像基准,用于评估模型在五个任务上的细粒度多模态推理能力:互补多模态思维链、多图像VQA、虚构图像描述、带幻觉触发的VQA,以及带有样本特定干扰项的细粒度检索。与现有基准不同,JourneyBench明确要求在异常虚构场景中进行细粒度多模态推理,仅靠语言偏见或整体图像概貌无法应对。我们在JourneyBench上对先进模型进行评估,并从多个细粒度维度分析其表现。所有五项任务的结果均显示,即使最先进模型也面临极大挑战,表明模型的视觉推理能力远未达到表面表现所示水平。我们讨论了研究意义,并提出未来方向。
原文摘要 · Abstract (English)
Existing vision-language understanding benchmarks largely consist of images of objects in their usual contexts. As a consequence, recent multimodal large language models can perform well with only a shallow visual understanding by relying on background language biases. Thus, strong performance on these benchmarks does not necessarily correlate with strong visual understanding. In this paper, we release JourneyBench, a comprehensive human-annotated benchmark of generated images designed to assess the model's fine-grained multimodal reasoning abilities across five tasks: complementary multimodal chain of thought, multi-image VQA, imaginary image captioning, VQA with hallucination triggers, and fine-grained retrieval with sample-specific distractors. Unlike existing benchmarks, JourneyBench explicitly requires fine-grained multimodal reasoning in unusual imaginary scenarios where language bias and holistic image gist are insufficient. We benchmark state-of-the-art models on JourneyBench and analyze performance along a number of fine-grained dimensions. Results across all five tasks show that JourneyBench is exceptionally challenging for even the best models, indicating that models' visual reasoning abilities are not as strong as they first appear. We discuss the implications of our findings and propose avenues for further research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。