首个面向科学图表的因果推理与事实验证基准,助力模型读懂数据可视化。
ClimateViz: A Benchmark for Statistical Reasoning and Fact Verification on Scientific Charts
- 构建4.9万条图表关联判断,含趋势、对比、因果等结构化解释
- 顶尖模型在图表理解上仅达77.8%准确率,远低于人类的92.7%
- 适合研究多模态推理、科学可信度评估的学者使用
科学事实核查长期聚焦文本与表格,忽视了呈现定量证据和统计推理的关键工具——科学图表。我们提出ClimateViz,首个基于专家标注图表的大规模科学事实核查基准。该数据集包含49,862条陈述与2,896个可视化图像,每条陈述被标注为支持、反驳或信息不足。为提升可解释性,每个样本均附带结构化知识图谱,涵盖趋势、比较与因果关系。我们在零样本与少样本设置下评估了当前最先进的多模态语言模型(包括闭源与开源系统)。结果表明,现有模型在图表推理任务中表现不佳:即使最佳系统如Gemini 2.5和InternVL 2.5,准确率也仅为76.2%至77.8%,远低于人类水平(89.3%和92.7%)。引入解释增强输出可部分提升部分模型性能。论文同时公开数据集与代码。
原文摘要 · Abstract (English)
Scientific fact-checking has mostly focused on text and tables, overlooking scientific charts, which are key for presenting quantitative evidence and statistical reasoning. We introduce ClimateViz, the first large-scale benchmark for scientific fact-checking using expert-curated scientific charts. ClimateViz contains 49,862 claims linked to 2,896 visualizations, each labeled as support, refute, or not enough information. To improve interpretability, each example includes structured knowledge graph explanations covering trends, comparisons, and causal relations. We evaluate state-of-the-art multimodal language models, including both proprietary and open-source systems, in zero-shot and few-shot settings. Results show that current models struggle with chart-based reasoning: even the best systems, such as Gemini 2.5 and InternVL 2.5, reach only 76.2 to 77.8 percent accuracy in label-only settings, far below human performance (89.3 and 92.7 percent). Explanation-augmented outputs improve performance in some models. We released our dataset and code alongside the paper.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。