arXiv:2505.18675cs.CVcs.AI2025-05被引 18

构建细粒度视觉推理新基准,揭示开源与闭源模型表现差异

ReasonMap: Towards Fine-Grained Visual Reasoning from Transit Maps

  • 设计跨城市高分辨率地铁图与1008组问答对评估视觉推理能力
  • 发现开源模型基础版比推理优化版更优,闭源模型则相反
  • 强调直接视觉定位比语言先验更重要,适合研究多模态推理的学者

多模态大语言模型在语义场景理解与图文对齐方面取得显著进展,其推理变体在涉及数学与逻辑的复杂任务中表现更佳。为填补这一差距,我们提出ReasonMap,一个专用于评估此类能力的新基准。该基准包含30个城市的高分辨率地铁图,涵盖1,008组问答对,覆盖两种问题类型和三种模板。我们设计了两级评估流程,以准确衡量答案正确性与质量。对16个主流MLLM的全面评估揭示了一个反直觉现象:在开源模型中,基础版本优于推理调优版本;而在闭源模型中则相反。在视觉掩码设置下的进一步分析表明,优异表现依赖于直接的视觉定位,而非仅依赖语言先验。我们还通过强化学习微调建立训练基线,为后续研究提供参考。我们希望本研究能为视觉推理提供新洞见,并帮助探究开源与闭源模型之间的差距。

原文摘要 · Abstract (English)

Multimodal large language models (MLLMs) have demonstrated significant progress in semantic scene understanding and text-image alignment, with reasoning variants enhancing performance on more complex tasks involving mathematics and logic. To bridge this gap, we introduce ReasonMap, a novel benchmark specifically designed to evaluate these capabilities. ReasonMap encompasses high-resolution transit maps from 30 cities and includes 1,008 question-answer pairs spanning two question types and three templates. Furthermore, we design a two-level evaluation pipeline that properly assesses answer correctness and quality. Our comprehensive evaluation of 16 popular MLLMs reveals a counterintuitive pattern: among open-source models, base variants outperform their reasoning-tuned counterparts, whereas the opposite trend is observed in closed-source models. Further analysis under the visual-masking setting confirms that strong performance necessitates direct visual grounding, rather than relying solely on language priors. We further establish a training baseline with reinforcement fine-tuning, providing a reference for future exploration. We hope this benchmark study offers new insights into visual reasoning and helps investigate the gap between open- and closed-source models.

视觉推理多模态地铁图基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。