arXiv:2511.10075cs.CL2025-11中稿 · AAAI被引 3

测试发现大模型看图表验证科学结论能力远不如看表格。

Format Matters: The Robustness of Multimodal LLMs in Reviewing Evidence from Tables and Charts

  • 用表格和图表双格式测试12个大模型的科学论断验证能力。
  • 模型在表格上准确率显著高于图表,差距达30%以上。
  • 小模型跨模态泛化差,适合需要图表理解的研究者关注。

随着投稿论文数量增长,辅助审稿人评估研究主张的系统需求日益增加。实验结果是科研工作的核心,常以表格或图表等形式呈现。当前多模态大语言模型在不同证据格式下验证科学主张的鲁棒性仍是一个重要且未被充分探索的问题。本文设计并开展一系列实验,评估多模态大模型利用表格和图表作为证据验证科学主张的能力。为此,我们对两个现有科学论文数据集进行改造,加入多模态主张验证所需的标注与结构。基于该适配数据集,我们评估了12个多模态大模型,发现当前模型在基于表格的证据上表现更好,而在基于图表的证据上表现较差。进一步的人类评估显示,人类在两种格式上均保持高绩效,而模型则表现不一。分析还发现,参数量小于80亿的小型多模态大模型在表格与图表任务间表现相关性弱,表明其跨模态泛化能力有限。这些发现凸显了当前模型在多模态推理能力上的关键差距。建议未来多模态大模型应更注重提升图表理解能力,以更好地支持科学主张验证。

原文摘要 · Abstract (English)

With the growing number of submitted scientific papers, there is an increasing demand for systems that can assist reviewers in evaluating research claims. Experimental results are a core component of scientific work, often presented in varying formats such as tables or charts. Understanding how robust current multimodal large language models (multimodal LLMs) are at verifying scientific claims across different evidence formats remains an important and underexplored challenge. In this paper, we design and conduct a series of experiments to assess the ability of multimodal LLMs to verify scientific claims using both tables and charts as evidence. To enable this evaluation, we adapt two existing datasets of scientific papers by incorporating annotations and structures necessary for a multimodal claim verification task. Using this adapted dataset, we evaluate 12 multimodal LLMs and find that current models perform better with table-based evidence while struggling with chart-based evidence. We further conduct human evaluations and observe that humans maintain strong performance across both formats, unlike the models. Our analysis also reveals that smaller multimodal LLMs (under 8B) show weak correlation in performance between table-based and chart-based tasks, indicating limited cross-modal generalization. These findings highlight a critical gap in current models' multimodal reasoning capabilities. We suggest that future multimodal LLMs should place greater emphasis on improving chart understanding to better support scientific claim verification.

多模态科学审稿图表理解

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。