用学术发表记录训练AI判断社科研究价值,效果优于专家平均水平。
LLMs learn scientific taste from institutional traces across the social sciences
- 用各学科发表记录作为奖励信号,微调大模型评估研究价值。
- 心理学模型准确率达85.5%,管理学模型比专家平均高17.6个百分点。
- 模型能自校准信心,高置信预测准确率极高,适合科研筛选场景。
强化学习推理已推动AI在可验证任务(如数学、代码)上的突破,但在低可验证性领域,缺乏明确评判标准,核心挑战在于判断哪些未验证的构想值得关注。本文测试机构痕迹——即各学科发表记录(内容、地点、层级)——能否作为AI评价者训练信号。覆盖八个社科领域(心理学、经济学、传播学、社会学、政治学、管理学、商学与金融、公共行政),构建了四层研究提案基准,并基于领域内发表结果对LLM进行监督微调。微调模型均超越25%随机基线,最佳单模型准确率从公共行政的55.0%到心理学的85.5%不等。在管理学中,与48位专家评审员、174名初级研究者及11个前沿推理模型对比,表现最优的Qwen3-4B模型达到59.2%准确率,较专家多数投票(41.6%,非平局)高出17.6个百分点,较前沿模型平均(31.1%)高出28.1个百分点。模型还表现出校准的信心:正确时信心上升,错误时下降,类似熟练审稿人“我确定”与“我猜测”的区分。仅基于高置信度预测的筛选,在各领域均达到极高准确率。结论表明,机构痕迹为科学评价中的低可验证判断提供了可扩展的训练信号。
原文摘要 · Abstract (English)
Reinforcement-learned reasoning has powered recent AI leaps on verifiable tasks, including mathematics, code, and structure prediction. The harder bottleneck is evaluative judgment in low-verifiability domains, where no oracle anchors reward and the core question is which untested ideas deserve attention. We test whether institutional traces, the record of what fields published, where, and at which tier, can serve as a training signal for AI evaluators. Across eight social science disciplines (psychology, economics, communication, sociology, political science, management, business and finance, public administration), we built held-out four-tier research-pitch benchmarks and supervised-fine-tuned (SFT) LLMs on field-specific publication outcomes. The fine-tuned models cleared the 25 percent chance baseline and exceeded frontier-model performance by wide margins, with best single-model accuracy ranging from 55.0 percent in public administration to 85.5 percent in psychology. In management, evaluated against 48 expert gatekeepers, 174 junior researchers, and 11 frontier reasoning models, the best single fine-tuned model (Qwen3-4B) reached 59.2 percent, 17.6 percentage points above expert majority vote (41.6 percent, non-tied) and 28.1 percentage points above the frontier mean (31.1 percent). The fine-tuned models also showed calibrated confidence: confidence rose when predictions were correct and fell when wrong, mirroring how a skilled reviewer can say "I'm sure" versus "I'm guessing." Selective triage on this signal reached very high accuracy on the highest-confidence subsets in every field. Institutional traces, we conclude, encode a scalable training signal for the low-verifiability judgment on which science depends.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。