构建大规模多领域图文一致性验证数据集,评测模型科学判断能力
M2-Verify: A Large-Scale Multidomain Benchmark for Checking Multimodal Claim Consistency

- 从医学与科研论文中提取46.9万条图文对,覆盖16个领域
- 顶尖模型在复杂图像变换下准确率降至61.6%,暴露出严重幻觉
- 专为评估科学结论与证据一致性设计,适合模型可信度研究者使用
评估科学论断需严格检验其与多模态证据的一致性。现有基准在规模、领域多样性和视觉复杂度上不足,难以真实评估这种对齐。为此,我们提出M2-Verify,一个大规模多模态数据集,用于检测科学论断的一致性。数据源自PubMed和arXiv,涵盖16个领域,共469,000余条实例,并经专家审核验证。大量基线实验表明,当前先进模型在保持一致性方面表现不佳:在低复杂度医学扰动下,顶级模型微平均F1最高达85.8%,但在高复杂度挑战(如解剖结构变化)下骤降至61.6%。此外,专家评估发现,模型生成的解释常出现幻觉。最后,我们展示了该数据集的实用性并提供完整使用指南。
原文摘要 · Abstract (English)
Evaluating scientific arguments requires assessing the strict consistency between a claim and its underlying multimodal evidence. However, existing benchmarks lack the scale, domain diversity, and visual complexity needed to evaluate this alignment realistically. To address this gap, we introduce M2-Verify, a large-scale multimodal dataset for checking scientific claim consistency. Sourced from PubMed and arXiv, M2-Verify provides over 469K instances across 16 domains, rigorously validated through expert audits. Extensive baseline experiments show that state-of-the-art models struggle to maintain robust consistency. While top models achieve up to 85.8\% Micro-F1 on low-complexity medical perturbations, performance drops to 61.6\% on high-complexity challenges like anatomical shifts. Furthermore, expert evaluations expose hallucinations when models generate scientific explanations for their alignment decisions. Finally, we demonstrate our dataset's utility and provide comprehensive usage guidelines.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。