评测大模型对论文论点的批判是否靠谱,发现仍远不如人类专家。
CLAIMCHECK: How Grounded are LLM Critiques of Scientific Papers?
- 构建标注数据集,分析模型对论文论点的批评是否准确对应
- 大模型在识别错误类型上表现尚可,但关联论点与批评仍不精准
- 适合研究AI评审、可信生成和科学验证的人看
科学同行评审的核心是针对论文提出的科学主张提供专业评价。尽管现在能自动生成看似合理(但泛化)的评审意见,但确保这些评价基于论文真实主张仍具挑战。为推动大模型在此类任务上的基准测试,我们引入CLAIMCHECK,这是一个从OpenReview中提取的NeurIPS 2023与2024投稿及其评审的标注数据集。该数据集由机器学习专家进行丰富标注,涵盖评审中的弱点陈述及其所质疑的论文主张,并包含弱点的有效性、客观性及类型等细粒度标签。我们在三个以论点为中心的任务上对多个大模型进行基准测试:(1) 将弱点与争议的论点关联;(2) 预测弱点的细粒度标签并改写以增强具体性;(3) 基于事实推理验证论文主张。实验表明,尽管前沿大模型在任务(2)中可预测弱点标签,但在其他任务上仍显著落后于人类专家。
原文摘要 · Abstract (English)
A core part of scientific peer review involves providing expert critiques that directly assess the scientific claims a paper makes. While it is now possible to automatically generate plausible (if generic) reviews, ensuring that these reviews are sound and grounded in the papers' claims remains challenging. To facilitate LLM benchmarking on these challenges, we introduce CLAIMCHECK, an annotated dataset of NeurIPS 2023 and 2024 submissions and reviews mined from OpenReview. CLAIMCHECK is richly annotated by ML experts for weakness statements in the reviews and the paper claims that they dispute, as well as fine-grained labels of the validity, objectivity, and type of the identified weaknesses. We benchmark several LLMs on three claim-centric tasks supported by CLAIMCHECK, requiring models to (1) associate weaknesses with the claims they dispute, (2) predict fine-grained labels for weaknesses and rewrite the weaknesses to enhance their specificity, and (3) verify a paper's claims with grounded reasoning. Our experiments reveal that cutting-edge LLMs, while capable of predicting weakness labels in (2), continue to underperform relative to human experts on all other tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。