arXiv:2603.26449cs.CL2026-03中稿 · NSLP@LREC 2026

构建气候事实核查与虚假叙事分类新基准,提升科学验证准确性。

ClimateCheck 2026: Scientific Fact-Checking and Disinformation Narrative Classification of Climate-related Claims

  • 融合密集检索与大模型推理,实现气候主张的自动验证。
  • 训练数据量翻三倍,新增虚假叙事分类任务,提升系统泛化能力。
  • 发现部分气候谎言难以验证,警示未来系统需针对性设计。

自动将气候相关主张与科学文献进行比对是一项挑战,因学术证据专业性强且气候虚假信息策略多样。ClimateCheck 2026 是该挑战的第二届共享任务,相比2025年版本,训练数据量增至三倍,并新增虚假叙事分类任务。比赛于2026年1月至2月在CodaBench平台举行,吸引20支参赛队伍、8次排行榜提交。参赛系统结合密集检索流水线、交叉编码器集成及具备结构化层次推理能力的大语言模型。除标准评估指标(Recall@K 和二元偏好)外,我们引入自动化框架,在标注不完整情况下评估检索质量,揭示传统指标排名中的系统性偏差。跨任务分析表明,并非所有气候虚假信息都同等可验证,暗示未来事实核查系统的设计应更精细化。

原文摘要 · Abstract (English)

Automatically verifying climate-related claims against scientific literature is a challenging task, complicated by the specialised nature of scholarly evidence and the diversity of rhetorical strategies underlying climate disinformation. ClimateCheck 2026 is the second iteration of a shared task addressing this challenge, expanding on the 2025 edition with tripled training data and a new disinformation narrative classification task. Running from January to February 2026 on the CodaBench platform, the competition attracted 20 registered participants and 8 leaderboard submissions, with systems combining dense retrieval pipelines, cross-encoder ensembles, and large language models with structured hierarchical reasoning. In addition to standard evaluation metrics (Recall@K and Binary Preference), we adapt an automated framework to assess retrieval quality under incomplete annotations, exposing systematic biases in how conventional metrics rank systems. A cross-task analysis further reveals that not all climate disinformation is equally verifiable, potentially implicating how future fact-checking systems should be designed.

气候核查虚假信息大模型评测基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。