测试大模型能否准确识别科学论文中论点与证据的逻辑关系。
Can AI Validate Science? Benchmarking LLMs for Accurate Scientific Claim $\rightarrow$ Evidence Reasoning
- 设计三轮提示法和逐条分析法提升模型匹配论点与证据的能力。
- 300多个跨领域的论点-证据对测试显示,闭源模型表现优于开源模型。
- 为大模型科学推理能力评估提供新基准,适合研究AI辅助科研者参考。
大语言模型在文献综述、创意生成和论文分析等复杂研究任务中应用日益广泛,但其对科学论文中论点与证据之间复杂逻辑关系的理解能力仍缺乏深入探索。本文提出CLAIM-BENCH,一个全面评估大模型在科学论点-证据提取与验证方面能力的基准,该任务反映对科学论证的深层理解。我们在六个不同领域共超过300个论点-证据对上系统比较了三种受分而治之启发的方法,涵盖三种闭源(GPT-4、Claude)与三种开源模型。结果表明,闭源模型在论点-证据识别任务中的精确率与召回率显著优于开源模型。通过设计三轮提示与逐条分析策略,可有效提升模型对分散证据与论点的关联能力,尽管计算成本更高。CLAIM-BENCH为评估大模型科学理解能力设立了新标准,既可作为诊断工具,也为构建能进行深度可靠推理的全篇论文分析系统指明方向。
原文摘要 · Abstract (English)
Large language models (LLMs) are increasingly being used for complex research tasks such as literature review, idea generation, and scientific paper analysis, yet their ability to truly understand and process the intricate relationships within complex research papers, such as the logical links between claims and supporting evidence remains largely unexplored. In this study, we present CLAIM-BENCH, a comprehensive benchmark for evaluating LLMs' capabilities in scientific claim-evidence extraction and validation, a task that reflects deeper comprehension of scientific argumentation. We systematically compare three approaches which are inspired by divide and conquer approaches, across six diverse LLMs, highlighting model-specific strengths and weaknesses in scientific comprehension. Through evaluation involving over 300 claim-evidence pairs across multiple research domains, we reveal significant limitations in LLMs' ability to process complex scientific content. Our results demonstrate that closed-source models like GPT-4 and Claude consistently outperform open-source counterparts in precision and recall across claim-evidence identification tasks. Furthermore, strategically designed three-pass and one-by-one prompting approaches significantly improve LLMs' abilities to accurately link dispersed evidence with claims, although this comes at increased computational cost. CLAIM-BENCH sets a new standard for evaluating scientific comprehension in LLMs, offering both a diagnostic tool and a path forward for building systems capable of deeper, more reliable reasoning across full-length papers.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。