arXiv:2505.15616cs.CV2025-05被引 7

构建多层级评测集,评估视觉语言模型从感知到推理的综合能力

LENS: Multi-level Evaluation of Multimodal Reasoning with Large Language Models

  • 设计三层任务体系:感知、理解、推理,覆盖12种日常场景
  • 包含3.4K张真实社交媒体图像,60万+人工标注问题,支持跨任务评估
  • 聚焦前沿模型表现,多数新模型在推理任务准确率不足60%

多模态大语言模型在融合视觉与语言信息方面取得显著进展,但在复杂现实场景中的推理能力仍受限。现有基准通常以任务为导向构建,不同任务样本未来自同一数据分布,难以评估低层感知能力对高层推理的协同作用。为此,我们提出LENS,一个包含3.4K张当代图像和60万+人工编写问题的多层级评测集,覆盖八个任务和十二种日常场景,分为感知、理解、推理三个递进层级。每张图像均配备所有任务的丰富标注,可支持模型在图像不变提示下完成从基础感知到组合推理的评测。图像源自社交媒体,其中53%发布于2025年1月之后。我们评估了15个以上前沿多模态模型,如Qwen2.5-VL-72B、InternVL3-78B、GPT-4o及两个推理模型QVQ-72B-preview和Kimi-VL,这些模型均发布于2024年12月之后,但在推理任务中准确率均未超过60%。

原文摘要 · Abstract (English)

Multimodal Large Language Models (MLLMs) have achieved significant advances in integrating visual and linguistic information, yet their ability to reason about complex and real-world scenarios remains limited. The existing benchmarks are usually constructed in the task-oriented manner without guarantee that different task samples come from the same data distribution, thus they often fall short in evaluating the synergistic effects of lower-level perceptual capabilities on higher-order reasoning. To lift this limitation, we contribute Lens, a multi-level benchmark with 3.4K contemporary images and 60K+ human-authored questions covering eight tasks and 12 daily scenarios, forming three progressive task tiers, i.e., perception, understanding, and reasoning. One feature is that each image is equipped with rich annotations for all tasks. Thus, this dataset intrinsically supports to evaluate MLLMs to handle image-invariable prompts, from basic perception to compositional reasoning. In addition, our images are manully collected from the social media, in which 53% were published later than Jan. 2025. We evaluate 15+ frontier MLLMs such as Qwen2.5-VL-72B, InternVL3-78B, GPT-4o and two reasoning models QVQ-72B-preview and Kimi-VL. These models are released later than Dec. 2024, and none of them achieve an accuracy greater than 60% in the reasoning tasks. Project page: https://github.com/Lens4MLLMs/lens. ICCV 2025 workshop page: https://lens4mllms.github.io/mars2-workshop-iccv2025/

多模态推理评测基准大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。