构建多模态科学结论验证数据集,测试模型对图表证据的推理能力。
MuSciClaims: Multimodal Scientific Claim Verification
- 从论文中自动提取并人工修改科学结论生成支持与矛盾命题
- 现有视觉语言模型验证准确率仅0.3-0.72,普遍高估支持度
- 暴露模型在定位证据、跨模态整合和理解图表基础结构上的缺陷
评估科学结论需识别、提取并推理信息丰富的图表中的多模态数据。尽管已有大量关于科学问答、图表描述及基于图表的多模态推理研究,但尚无直接测试结论验证能力的现成多模态基准。为此,我们引入新基准MuSciClaims,并配套诊断任务。我们自动从科学论文中提取支持性结论,并人工扰动生成矛盾结论,扰动设计用于测试特定验证能力。同时引入一系列诊断任务以分析模型失败原因。结果表明,多数视觉语言模型表现较差(F1值0.3–0.5),即使最优模型也仅达0.72。模型普遍倾向于判断结论为支持,可能误读结论中的细微扰动。诊断显示,模型在定位正确证据、跨模态信息聚合及理解图表基本构成方面存在显著不足。
原文摘要 · Abstract (English)
Assessing scientific claims requires identifying, extracting, and reasoning with multimodal data expressed in information-rich figures in scientific literature. Despite the large body of work in scientific QA, figure captioning, and other multimodal reasoning tasks over chart-based data, there are no readily usable multimodal benchmarks that directly test claim verification abilities. To remedy this gap, we introduce a new benchmark MuSciClaims accompanied by diagnostics tasks. We automatically extract supported claims from scientific articles, which we manually perturb to produce contradicted claims. The perturbations are designed to test for a specific set of claim verification capabilities. We also introduce a suite of diagnostic tasks that help understand model failures. Our results show most vision-language models are poor (~0.3-0.5 F1), with even the best model only achieving 0.72 F1. They are also biased towards judging claims as supported, likely misunderstanding nuanced perturbations within the claims. Our diagnostics show models are bad at localizing correct evidence within figures, struggle with aggregating information across modalities, and often fail to understand basic components of the figure.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。