用真实气候问题评估大模型,发现其会编造证据。
CLINB: A Climate Intelligence Benchmark for Foundational Models
- 基于真实用户提问和科学家制定的评分标准构建评测集。
- 前沿模型知识水平达博士级,但引用文献和图片多为虚构。
- 适合关注可信AI、科学问答与模型幻觉的研究者使用。
评估大型语言模型(LLMs)处理复杂专业知识的能力仍是关键挑战。我们以气候变化为视角,提出CLINB基准,用于评估模型在开放性、有依据、多模态问答任务中的表现,要求知识质量与证据支持清晰可靠。CLINB基于真实用户问题数据集及由顶尖气候科学家制定的评估细则。我们实施并验证了基于模型的评估流程,测试了多个前沿模型。结果揭示显著矛盾:前沿模型具备出色的跨知识融合能力,常表现出博士级的理解与表达水平,优于由领域专家结合弱模型生成的‘混合’答案;但其在事实锚定方面存在严重缺陷,证据质量参差不齐,引用文献和图像的幻觉率较高。我们认为,弥合知识整合与可验证溯源之间的差距,是实现AI在科研流程中可信部署的关键,而像CLINB这样可靠且可解释的基准对构建可信AI系统至关重要。
原文摘要 · Abstract (English)
Evaluating how Large Language Models (LLMs) handle complex, specialized knowledge remains a critical challenge. We address this through the lens of climate change by introducing CLINB, a benchmark that assesses models on open-ended, grounded, multimodal question answering tasks with clear requirements for knowledge quality and evidential support. CLINB relies on a dataset of real users' questions and evaluation rubrics curated by leading climate scientists. We implement and validate a model-based evaluation process and evaluate several frontier models. Our findings reveal a critical dichotomy. Frontier models demonstrate remarkable knowledge synthesis capabilities, often exhibiting PhD-level understanding and presentation quality. They outperform "hybrid" answers curated by domain experts assisted by weaker models. However, this performance is countered by failures in grounding. The quality of evidence varies, with substantial hallucination rates for references and images. We argue that bridging this gap between knowledge synthesis and verifiable attribution is essential for the deployment of AI in scientific workflows and that reliable, interpretable benchmarks like CLINB are needed to progress towards building trustworthy AI systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。