arXiv:2601.19925cs.CLcs.AI2026-01被引 1

测试三款大模型评估学术摘要的能力,发现其对客观内容判断较准。

Evaluating Large Language Models for Abstract Evaluation Tasks: An Empirical Study

  • 用统一评分标准让人类与三款大模型评分160篇摘要
  • 大模型间一致性良好,与人类在客观指标上中等吻合
  • 主观评价仍需人工把关,适合批量初筛场景

大型语言模型(LLMs)能处理请求并生成文本,但其在评估复杂学术内容方面的可行性有待深入探讨。为探究大模型在辅助科学评审中的潜力,本研究比较了ChatGPT-5、Gemini-3-Pro和Claude-Sonnet-4.5在评估摘要时的一致性与可靠性,对比对象包括其他模型及人类评审员。从本地会议选取160篇摘要,由14名人类评审员和三款大模型按同一评分标准打分。分析了三款大模型与十四名评审员的综合得分分布,计算组内信度(组内ICC)和模型-人类一致性。使用Bland-Altman图观察共识模式与系统性偏差。结果显示,大模型之间达成良好至优秀的一致性(ICC:0.59–0.87)。ChatGPT与Claude在整体质量与内容特定维度上与人类评审员达到中等一致性,复合分、印象、清晰度、目标、结果等维度的ICC约为0.45–0.60;在主观维度上表现中等,影响、参与度、适用性等的ICC范围为0.23–0.38。Gemini在半数标准上达到中等一致性,但在影响与适用性上无可靠性。三款模型与人类平均分的差异均在可接受范围内(ChatGPT=0.24,Gemini=0.42,Claude=-0.02)。讨论表明,大模型可在批量处理摘要时,以中等程度与人类专家一致地评估整体质量和客观标准。通过合理架构设计,其可在人力难以应对的海量稿件评审中发挥作用。然而在主观维度上的较弱表现表明,人工智能应作为补充角色,人类专业判断仍不可或缺。

原文摘要 · Abstract (English)

Introduction: Large language models (LLMs) can process requests and generate texts, but their feasibility for assessing complex academic content needs further investigation. To explore LLM's potential in assisting scientific review, this study examined ChatGPT-5, Gemini-3-Pro, and Claude-Sonnet-4.5's consistency and reliability in evaluating abstracts compared to one another and to human reviewers. Methods: 160 abstracts from a local conference were graded by human reviewers and three LLMs using one rubric. Composite score distributions across three LLMs and fourteen reviewers were examined. Inter-rater reliability was calculated using intraclass correlation coefficients (ICCs) for within-AI reliability and AI-human concordance. Bland-Altman plots were examined for visual agreement patterns and systematic bias. Results: LLMs achieved good-to-excellent agreement with each other (ICCs: 0.59-0.87). ChatGPT and Claude reached moderate agreement with human reviewers on overall quality and content-specific criteria, with ICCs ~.45-.60 for composite, impression, clarity, objective, and results. They exhibited fair agreement on subjective dimensions, with ICC ranging from 0.23-0.38 for impact, engagement, and applicability. Gemini showed fair agreement on half criteria and no reliability on impact and applicability. Three LLMs showed acceptable or negligible mean difference (ChatGPT=0.24, Gemini=0.42, Claude=-0.02) from the human mean composite scores. Discussion: LLMs could process abstracts in batches with moderate agreement with human experts on overall quality and objective criteria. With appropriate process architecture, they can apply a rubric consistently across volumes of abstracts exceeding feasibility for a human rater. The weaker performance on subjective dimensions indicates that AI should serve a complementary role in evaluation, while human expertise remains essential.

大模型评估学术评审人机协作

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。