首个评估大模型科学论断多模态验证能力的基准测试
SciVer: Evaluating Foundation Models for Multimodal Scientific Claim Verification
- 构建3000个专家标注的多模态科学论断验证样本
- 21个主流多模态模型在该基准上表现远低于人类专家
- 揭示开源模型在科学文献理解与推理中的关键缺陷
我们提出SciVer,首个专门用于评估基础模型在多模态科学语境下验证论断能力的基准。SciVer包含来自1,113篇科学论文的3,000个专家标注样本,涵盖四种常见推理类型。每个样本均附有专家标注的支持证据,以实现细粒度评估。我们评估了21个前沿多模态基础模型(包括o4-mini、Gemini-2.5-Flash、Llama-3.2-Vision和Qwen2.5-VL)的表现。实验显示这些模型在SciVer上的表现与人类专家存在显著差距。通过检索增强生成(RAG)分析及人工错误评估,我们识别出当前开源模型在理解与推理多模态科学文献中的关键局限,为提升模型能力提供重要洞见。
原文摘要 · Abstract (English)
We introduce SciVer, the first benchmark specifically designed to evaluate the ability of foundation models to verify claims within a multimodal scientific context. SciVer consists of 3,000 expert-annotated examples over 1,113 scientific papers, covering four subsets, each representing a common reasoning type in multimodal scientific claim verification. To enable fine-grained evaluation, each example includes expert-annotated supporting evidence. We assess the performance of 21 state-of-the-art multimodal foundation models, including o4-mini, Gemini-2.5-Flash, Llama-3.2-Vision, and Qwen2.5-VL. Our experiment reveals a substantial performance gap between these models and human experts on SciVer. Through an in-depth analysis of retrieval-augmented generation (RAG), and human-conducted error evaluations, we identify critical limitations in current open-source models, offering key insights to advance models' comprehension and reasoning in multimodal scientific literature tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。