arXiv:2510.22389cs.DLcs.AI2025-10被引 18

小模型也能准确评估论文质量,平均打分效果最好。

Can Small and Reasoning Large Language Models Score Journal Articles for Research Quality and Do Averaging and Few-shot Help?

  • 用多个相同问题重复评分,显著提升小模型的评估能力。
  • 40亿参数以上模型可稳定评高质量,10亿参数常不足。
  • 推理模式虽慢但无优势,适合资源受限场景部署。

先前研究显示,基于云的大型语言模型(如ChatGPT、Gemini)和中等规模开源模型Gemma3 27b在期刊论文质量评分上与专家评分中度相关。本文评估其他中等规模、小型及推理型模型在此任务中的表现,使用包含2,780篇医学、健康与生命科学领域论文的数据集,涵盖6个学科,并采用两种黄金标准(其中一种为新提出)。测试了少样本提示与评分平均化策略。结果表明,中等规模模型性能接近ChatGPT 4o-mini和Gemini 2.0 Flash;10亿参数模型通常不足,40亿参数模型有时也不够。推理模型未表现出明显优势。此外,多次相同查询的评分平均化是普遍有效的策略,少量示例(四例)可能略有帮助。首次证明,超过40亿参数的小型模型具备较强论文质量评估能力,尤其配合评分平均化时;而推理模式因速度慢且无增益,不推荐使用。这使利用模型辅助科研评价更具可信度,多种模型可离线部署,无需大量计算资源。

原文摘要 · Abstract (English)

Previous research has shown that journal article quality ratings from the cloud based Large Language Model (LLM) families ChatGPT and Gemini and the medium sized open weights LLM Gemma3 27b correlate moderately with expert research quality scores. This article assesses whether other medium sized LLMs, smaller LLMs, and reasoning models have similar abilities. This is tested with Gemma3 variants, Llama4 Scout, Qwen3, Magistral Small and DeepSeek R1 on a dataset of 2,780 medical, health and life science papers in 6 fields, with two different gold standards, one novel. Few-shot and score averaging approaches are also evaluated. The results suggest that medium-sized LLMs have similar performance to ChatGPT 4o-mini and Gemini 2.0 Flash, but that 1b parameters may often, and 4b sometimes, be too few. Reasoning models did not have a clear advantage. Moreover, averaging scores from multiple identical queries seems to be a universally successful strategy, and there is weak evidence that few-shot prompts (four examples) tend to help. Overall, the results show, for the first time, that smaller LLMs >4b have a substantial capability to rate journal articles for research quality, especially if score averaging is used, but that reasoning does not give an advantage for this task; it is therefore not recommended because it is slow. The use of LLMs to support research evaluation is now more credible since multiple variants have a similar ability, including many that can be deployed offline in a secure environment without substantial computing resources.

论文评估小模型评分平均离线部署

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。