arXiv:2508.07196cs.DLcs.AI2025-08被引 12

小模型也能准确评估科研质量,适合离线使用。

Can Smaller Large Language Models Evaluate Research Quality?

  • 用270亿参数的Gemma-3-27b-it模型评估10万篇论文质量
  • 相关性达ChatGPT 4o的83.8%,4o-mini的94.7%
  • 无需重复运行,适合低成本、安全离线场景

尽管Google Gemini(1.5 Flash)和ChatGPT(4o与4o-mini)在几乎所有领域中,其科研质量评分与专家评分呈正相关,且与引用次数的相关性更强,但小型大语言模型(LLM)是否具备此能力尚不明确。为此,本文评估了可下载的60GB Gemma-3-27b-it模型。在对104,187篇论文的评估中,该模型在英国研究卓越框架2021年34个学科单位中的所有领域,其评分与专家质量代理指标均呈正相关。其相关性强度达到ChatGPT 4o的83.8%和4o-mini的94.7%。与两个更大模型不同的是,该模型在五次重复平均后相关性未显著提升,其评分整体偏低,报告风格相对统一。结果表明,科研质量评估可在本地部署的模型上完成,这一能力并非仅存在于最大模型中。此外,通过重复提升评分并非所有模型的普遍特性。结论:尽管最大模型仍具最优能力,但小型模型亦可胜任科研评估任务,在降低成本或需安全离线处理时尤为有用。

原文摘要 · Abstract (English)

Although both Google Gemini (1.5 Flash) and ChatGPT (4o and 4o-mini) give research quality evaluation scores that correlate positively with expert scores in nearly all fields, and more strongly that citations in most, it is not known whether this is true for smaller Large Language Models (LLMs). In response, this article assesses Google's Gemma-3-27b-it, a downloadable LLM (60Gb). The results for 104,187 articles show that Gemma-3-27b-it scores correlate positively with an expert research quality score proxy for all 34 Units of Assessment (broad fields) from the UK Research Excellence Framework 2021. The Gemma-3-27b-it correlations have 83.8% of the strength of ChatGPT 4o and 94.7% of the strength of ChatGPT 4o-mini correlations. Differently from the two larger LLMs, the Gemma-3-27b-it correlations do not increase substantially when the scores are averaged across five repetitions, its scores tend to be lower, and its reports are relatively uniform in style. Overall, the results show that research quality score estimation can be conducted by offline LLMs, so this capability is not an emergent property of the largest LLMs. Moreover, score improvement through repetition is not a universal feature of LLMs. In conclusion, although the largest LLMs still have the highest research evaluation score estimation capability, smaller ones can also be used for this task, and this can be helpful for cost saving or when secure offline processing is needed.

科研评估小模型离线部署Gemma

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。