arXiv:2512.10895cs.AI2025-12

用大模型辅助大型科研设施提案筛选,更高效公正。

LLMs Can Assist with Proposal Selection at Large User Facilities

  • 用大模型对比提案优劣,解决人工评审不一致问题
  • 模型排名与人工评分相关性达0.2-0.8,剔除10%异常后超0.5
  • 成本降低超90%,还能分析提案相似度等深层信息

我们研究大语言模型(LLMs)在大型用户设施提案筛选中的应用,提供一种可扩展、一致且低成本的替代方案。传统人工评审存在提案间相关性弱、主观偏见和不一致等问题。基于成对偏好排序的方法理论上更优,但其二次复杂度使人工难以实施。本文利用美国橡树岭国家实验室散裂中子源(SNS)三个光束线的高质量提案与发表记录,证明大模型排名与人工排名高度相关(斯皮尔曼相关系数ρ≈0.2–0.8,剔除10%异常值后≥0.5)。此外,大模型在识别高发表潜力提案方面表现不逊于人类,且成本降低两个数量级以上。除排名外,大模型还支持嵌入模型实现提案相似性量化分析,为评审委员会提供关键决策信息。

原文摘要 · Abstract (English)

We explore how large language models (LLMs) can enhance the proposal selection process at large user facilities, offering a scalable, consistent, and cost-effective alternative to traditional human review. Proposal selection depends on assessing the relative strength among submitted proposals; however, traditional human scoring often suffers from weak inter-proposal correlations and is subject to reviewer bias and inconsistency. A pairwise preference-based approach is logically superior, providing a more rigorous and internally consistent basis for ranking, but its quadratic workload makes it impractical for human reviewers. We address this limitation using LLMs. Leveraging the uniquely well-curated proposals and publication records from three beamlines at the Spallation Neutron Source (SNS), Oak Ridge National Laboratory (ORNL), we show that the LLM rankings correlate strongly with the human rankings (Spearman $ρ\simeq 0.2-0.8$, improving to $\geq 0.5$ after 10\% outlier removal). Moreover, LLM performance is no worse than that of human reviewers in identifying proposals with high publication potential, while costing over two orders of magnitude less. Beyond ranking, LLMs enable advanced analyses that are challenging for humans, such as quantitative assessment of proposal similarity via embedding models, which provides information crucial for review committees.

大模型应用科研管理智能评审推荐系统

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。