测试大模型对科研资助申请书的评审能力,发现分段分析效果最好。
Evaluating LLM-Based Grant Proposal Review via Structured Perturbations
- 通过六类结构化扰动,测试大模型在六个维度上的敏感性。
- 分段评审检测率和评分可靠性显著优于其他方法,但整体波动大。
- 适合关注评审自动化与模型局限性的研究管理者或审稿人。
随着AI辅助科研申请的数量超过人工评审能力,形成科研生态中的‘马尔萨斯陷阱’,本文探究大模型在高风险评审任务中的能力与局限。基于六份英国工程与物理科学研究理事会(EPSRC)的申请书,构建一种基于扰动的评估框架,探测大模型在经费、时间表、能力、契合度、清晰度和影响力六个质量维度上的敏感性。对比三种评审架构:单次评审、逐段分析、以及模拟专家小组的‘人物议会’集成方法。结果表明,逐段分析法在检测率和评分可靠性上显著优于其他方法,而计算成本高的议会方法表现不优于基线。不同扰动类型检测效果差异显著,契合度问题易被识别,但清晰度缺陷普遍被遗漏。人工评估显示,大模型反馈基本有效,但更偏向合规检查而非整体评估。结论认为,当前大模型可在EPSRC评审中提供辅助价值,但存在高波动性和评审优先级错配问题。代码与非受保护数据已公开。
原文摘要 · Abstract (English)
As AI-assisted grant proposals outpace manual review capacity in a kind of ``Malthusian trap'' for the research ecosystem, this paper investigates the capabilities and limitations of LLM-based grant reviewing for high-stakes evaluation. Using six EPSRC proposals, we develop a perturbation-based framework probing LLM sensitivity across six quality axes: funding, timeline, competency, alignment, clarity, and impact. We compare three review architectures: single-pass review, section-by-section analysis, and a 'Council of Personas' ensemble emulating expert panels. The section-level approach significantly outperforms alternatives in both detection rate and scoring reliability, while the computationally expensive council method performs no better than baseline. Detection varies substantially by perturbation type, with alignment issues readily identified but clarity flaws largely missed by all systems. Human evaluation shows LLM feedback is largely valid but skewed toward compliance checking over holistic assessment. We conclude that current LLMs may provide supplementary value within EPSRC review but exhibit high variability and misaligned review priorities. We release our code and any non-protected data.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。