构建跨模态科学论断验证数据集,含真实被证伪的论断。
SciClaimEval: Cross-modal Claim Verification in Scientific Papers
- 从真实论文中提取论断,通过修改图表证据生成反例。
- 包含1664个样本,覆盖三大领域,经专家标注验证。
- 揭示当前模型在图像驱动验证上仍远落后于人类。
我们提出SciClaimEval,一个用于科学论断验证的新数据集。与现有资源不同,该数据集包含直接从已发表论文中提取的真实论断,包括被证伪的论断。为生成被证伪的论断,我们引入一种新方法:修改支持性证据(图表),而非改动论断本身或依赖大语言模型生成矛盾内容。数据集提供跨模态证据,涵盖多种表现形式:图像以图片形式呈现,表格则以图像、LaTeX源码、HTML和JSON等多种格式提供。SciClaimEval共包含180篇论文中的1664个标注样本,覆盖机器学习、自然语言处理和医学三个领域,并经过专家验证。我们对11个多模态基础模型(开源与专有)进行了基准测试。结果表明,基于图像的验证对所有模型而言仍是重大挑战,最优系统与人类基线之间仍存在显著性能差距。
原文摘要 · Abstract (English)
We present SciClaimEval, a new scientific dataset for the claim verification task. Unlike existing resources, SciClaimEval features authentic claims, including refuted ones, directly extracted from published papers. To create refuted claims, we introduce a novel approach that modifies the supporting evidence (figures and tables), rather than altering the claims or relying on large language models (LLMs) to fabricate contradictions. The dataset provides cross-modal evidence with diverse representations: figures are available as images, while tables are provided in multiple formats, including images, LaTeX source, HTML, and JSON. SciClaimEval contains 1,664 annotated samples from 180 papers across three domains, machine learning, natural language processing, and medicine, validated through expert annotation. We benchmark 11 multimodal foundation models, both open-source and proprietary, across the dataset. Results show that figure-based verification remains particularly challenging for all models, as a substantial performance gap remains between the best system and human baseline.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。