arXiv:2605.18752cs.IRastro-ph.IM2026-05

传统统计方法比生成式AI更准地找出领域专家。

Traditional statistical representations outperform generative AI in identifying expert peer reviewers

论文配图:Traditional statistical representations outperform generative AI in identifying expert peer reviewers
图 1 · 摘自论文原文
  • 用词频逆文档频率法识别专家,效果优于大模型。
  • 传统方法在前25名推荐中命中专家率达79.5%。
  • 适合需要精准细分领域匹配的科研评审场景。

科学论文投稿量激增已使同行评审系统承压。尽管研究人员规模扩大,但手动筛选专家已不可行,机构转向使用大型语言模型(LLMs)自动化专家识别。然而,这些模型在准确识别领域专家方面的可靠性尚未经过严格评估。我们对统计与AI驱动的专家识别方法进行了全面实证评估,将专家识别视为信息检索问题,并以一个主要国际天文台的分布式评审系统为数据源,以提案作者身份作为领域专长的代理真实标签。评估了六种观测站与计算机科学会议中使用的方法,结果表明传统统计表示优于生成式AI:词频逆文档频率(TF-IDF)在前25名推荐中成功识别出标注专家的比例达79.5%,而GPT-4o mini仅为51.5%。研究指出,区分子领域专长需精细词汇表达,而生成式方法的语义平滑会掩盖这种细微差异。本研究建立了一个严谨的自动化同行评审评估框架,证明透明可复现的统计方法在专业科学任务中仍优于计算成本高昂的大模型。

原文摘要 · Abstract (English)

The exponential growth of scientific submissions has strained the peer review system. Despite the rapidly expanding global pool of researchers, this unprecedented scale has rendered the previous approach of manual expert identification unfeasible. Therefore, institutions have naturally turned to Large Language Models (LLMs) to automate intricate processes like expert reviewer identification. However, the reliability of these new models in accurately identifying domain experts lacks rigorous evaluation. We conduct a comprehensive empirical evaluation of statistical and AI-driven expertise identification methodologies to benchmark their reliability and limitations. Framing expert identification as an information retrieval problem, we utilize the distributed peer review system of a major international astronomical observatory, where proposal authorship serves as our proxy ground truth for domain expertise. Evaluating six retrieval methodologies utilized across observatories and computer science conferences, we demonstrate that traditional statistical representations outperform generative AI. Specifically, Term Frequency-Inverse Document Frequency successfully identified a labeled expert within the top 25 recommendations 79.5% of the time, compared to 51.5% for GPT-4o mini. Our results highlight that distinguishing subfield expertise requires fine-grained vocabulary, which is obscured by the semantic smoothing in generative methods. By establishing a rigorous evaluation framework for automated peer review, we demonstrate that transparent and reproducible statistical representations still outperform computationally expensive LLMs in specialized scientific tasks.

同行评审专家识别统计方法大模型评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。