用大模型批量评1200份本科生研究申请,4小时完成评分与理由生成。
Using Large Language Models to Support High Volume Application Review for an Undergraduate Research Program

- 用GPT-5.2模型按六项标准打分,每篇平均14秒处理
- 模型输出评分与理由,与人工评分差异在低分段更明显
- 让项目负责人4小时完成筛选,比过去节省数周时间
普渡大学的暑期本科生研究奖学金(SURF)每年接收数千份申请,评估工作量巨大。本文介绍基于大语言模型(LLM)的辅助评审工具开发与初步部署,用于处理2026届SURF计划约1,200份个人陈述(SoP)。采用OpenAI GPT系列模型(GPT-4o、GPT-5-mini、GPT-5.2),依据六项子维度的结构化评分表(每项0-3分)进行评分。通过少量人工标注样本微调模型输出。使用GPT-5.2模型处理全部1,200份文档耗时约4.6小时,平均每篇14秒(篇幅500-2,000词)。模型对评分标准的遵循度以GPT-5.2最高;低分段评分分歧较大。模型输出包含数值评分、评语及原文摘录,替代了以往分布式人工评分角色。项目协调人结合模型输出与原始申请材料,参照往年标准完成最终候选名单筛选,耗时约4小时,显著优于以往多周的协调周期。
原文摘要 · Abstract (English)
Undergraduate research programs such as the Summer Undergraduate Research Fellowship (SURF) at Purdue University receive thousands of applications every year, requiring significant time and effort for program staff to evaluate each submission consistently and within tight timelines. This work-in-progress paper describes the development and initial deployment of a large language model (LLM)-based tool to assist in the evaluation of approximately 1,200 student Statements of Purpose (SoPs) for the SURF 2026 cycle at Purdue University. The workflow utilizes OpenAI GPT models (GPT-4o, GPT-5-mini, and GPT-5.2) and uses a structured rubric across six subcategories, each scored on a 0-3 scale. A few SoPs, graded by program staff, were used to tune the model responses. The model prompt was designed to generate both numerical scores, rationales (including positive and negative aspects) and short excerpts from each submission. Using GPT-5.2, the full batch of 1,200 SoPs was processed in approximately 4.6 hours of compute time, averaging roughly 14 seconds per SoP (with per-SoP timing varying with SoP length, which ranged from 500 to 2,000 words). Notable differences in rubric adherence were observed across model versions, with GPT-5.2 adhering most closely. Disagreement in model scores was more pronounced for lower-scoring submissions. The LLM outputs replicated the role previously played by distributed human graders, providing the program coordinator with scored and rationale-annotated outputs for the entire applicant pool. The program coordinator then reviewed these outputs alongside each applicant's SoP, applying the same downstream office criteria used in prior SURF cycles, to produce a shortlist of strong candidates. This coordinator review was completed in approximately 4 hours, compared to the multi-week coordination effort required in prior program cycles.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。