用大模型评估论文质量,粗粒度效果好,细粒度仍有差距。
Large language models for post-publication research evaluation: Evidence from expert recommendations and citation indicators
- 对比专家意见和引用数据,测试大模型在论文评价中的表现。
- 粗粒度识别优质论文准确率超0.8,细粒度评分效果下降明显。
- 微调+少样本提示效果最佳,但与引用指标相关性中等。
科学论文质量评估对学术交流至关重要,但现有方法存在可扩展性差、主观性强和延迟高问题。大语言模型(LLMs)为基于文本内容的自动化评估带来新机遇。本研究通过对比专家推荐与引用指标,检验LLMs在发表后同行评审任务中的适用性。基于H1 Connect平台文章构建两项任务:识别高质量论文,以及更细粒度的评分、重要性分类和专家风格评论。测试了BERT模型、通用LLM及推理导向型模型,采用多种学习策略。结果表明,LLMs在粗粒度任务中表现良好,识别高推荐论文准确率超过0.8;但在细粒度评分任务中性能显著下降。少样本提示优于零样本,而监督微调效果最强且最均衡。检索增强提示在部分情况下提升分类准确率,但未持续增强与引用指标的一致性。模型输出与引用指标总体呈正相关,但相关性仅为中等。
原文摘要 · Abstract (English)
Assessing the quality of scientific research is essential for scholarly communication, yet widely used approaches face limitations in scalability, subjectivity, and time delay. Recent advances in large language models (LLMs) offer new opportunities for automated research evaluation based on textual content. This study examines whether LLMs can support post-publication peer review tasks by benchmarking their outputs against expert judgments and citation-based indicators. Two evaluation tasks are constructed using articles from the H1 Connect platform: identifying high-quality articles and performing finer-grained evaluation including article rating, merit classification, and expert style commenting. Multiple model families, including BERT models, general-purpose LLMs, and reasoning oriented LLMs, are evaluated under multiple learning strategies. Results show that LLMs perform well in coarse grained evaluation tasks, achieving accuracy above 0.8 in identifying highly recommended articles. However, performance decreases substantially in fine-grained rating tasks. Few-shot prompting improves performance over zero-shot settings, while supervised fine-tuning produces the strongest and most balanced results. Retrieval augmented prompting improves classification accuracy in some cases but does not consistently strengthen alignment with citation indicators. The overall correlations between model outputs and citation indicators remain positive but moderate.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。