arXiv:2503.16856cs.CL2025-03ICCV被引 6

评测视觉语言模型跨源推理能力,发现当前模型表现远未达标。

MMCR: Benchmarking Cross-Source Reasoning in Scientific Papers

  • 构建276道跨源科学论文推理题,覆盖7个学科10类任务。
  • 顶尖模型GPT-4o仅达48.55%准确率,多表理解任务仅20%。
  • 大模型用思维链提升效果,小模型反而更差,提示训练策略需优化。

机器全面理解科学论文体现了高级通用人工智能水平,要求在碎片化、异构信息源间进行推理,具有高度复杂性和现实意义。尽管视觉语言模型(VLMs)在单图像或单文本推理任务中取得显著进展,但其利用跨源信息进行推理的能力仍待解决。本文提出MMCR,一个高难度基准,用于评估VLM在科学论文中的跨源推理能力。该基准包含276道由人工精心标注的高质量问题,覆盖7个学科和10种任务类型。对18个VLM的实验表明,跨源推理对现有模型构成巨大挑战:即使表现最佳的GPT-4o也仅达到48.55%的整体准确率,多表理解任务准确率仅为20%;次优模型Qwen2.5-VL-72B整体准确率为39.86%。此外,我们研究了思维链(CoT)技术的影响,发现其对小模型产生负面影响,而大模型则表现出显著性能提升。这些结果凸显了发展有效利用跨源信息进行推理的VLM的紧迫性。

原文摘要 · Abstract (English)

Fully comprehending scientific papers by machines reflects a high level of Artificial General Intelligence, requiring the ability to reason across fragmented and heterogeneous sources of information, presenting a complex and practically significant challenge. While Vision-Language Models (VLMs) have made remarkable strides in various tasks, particularly those involving reasoning with evidence source from single image or text page, their ability to use cross-source information for reasoning remains an open problem. This work presents MMCR, a high-difficulty benchmark designed to evaluate VLMs' capacity for reasoning with cross-source information from scientific papers. The benchmark comprises 276 high-quality questions, meticulously annotated by humans across 7 subjects and 10 task types. Experiments with 18 VLMs demonstrate that cross-source reasoning presents a substantial challenge for existing models. Notably, even the top-performing model, GPT-4o, achieved only 48.55% overall accuracy, with only 20% accuracy in multi-table comprehension tasks, while the second-best model, Qwen2.5-VL-72B, reached 39.86% overall accuracy. Furthermore, we investigated the impact of the Chain-of-Thought (CoT) technique on cross-source reasoning and observed a detrimental effect on small models, whereas larger models demonstrated substantially enhanced performance. These results highlight the pressing need to develop VLMs capable of effectively utilizing cross-source information for reasoning.

跨源推理科学论文视觉语言模型基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。