通过优化示例边界提升大模型自动评分的准确性与一致性。
Optimizing In-Context Demonstrations for LLM-based Automated Grading
- 将示例选择重构为边界聚焦的优化问题,识别相似但得分不同的样本对。
- 生成区分性评语,使模型更精准把握评分标准边缘,显著提升临界案例表现。
- 适用于需要高一致性和可解释性的教育自动化评估场景。
开放性学生作答的自动化评估是实现个性化反馈规模化的重要能力。尽管大语言模型(LLMs)通过上下文学习(ICL)在评分任务中展现潜力,其可靠性高度依赖于少样本示例的选择和高质量评语的构建。传统检索方法通常基于语义相似性选取示例,难以捕捉评分量表中细微的决策边界。此外,手动设计引导模型的专家评语也构成显著瓶颈。为此,我们提出GUIDE(Grading Using Iteratively Designed Exemplars)框架,将示例选择与优化重构为边界聚焦的优化问题。GUIDE采用连续的选样与精炼循环,引入新颖的对比操作符,识别语义相近但得分不同的“边界对”。通过生成明确说明为何得某分而非相邻分数的判别性评语来增强示例。在物理、化学及教学知识数据集上的大量实验表明,GUIDE显著优于标准检索基线。该方法聚焦评分量表边缘,对临界案例表现出极强鲁棒性,且提升量表遵循度。GUIDE为建立贴近人类教学标准、可信可扩展的评估系统开辟了新路径。
原文摘要 · Abstract (English)
Automated assessment of open-ended student responses is a critical capability for scaling personalized feedback in education. While large language models (LLMs) have shown promise in grading tasks via in-context learning (ICL), their reliability is heavily dependent on the selection of few-shot exemplars and the construction of high-quality rationales. Standard retrieval methods typically select examples based on semantic similarity, which often fails to capture subtle decision boundaries required for rubric adherence. Furthermore, manually crafting the expert rationales needed to guide these models can be a significant bottleneck. To address these limitations, we introduce GUIDE (Grading Using Iteratively Designed Exemplars), a framework that reframes exemplar selection and refinement in automated grading as a boundary-focused optimization problem. GUIDE operates on a continuous loop of selection and refinement, employing novel contrastive operators to identify "boundary pairs" that are semantically similar but possess different grades. We enhance exemplars by generating discriminative rationales that explicitly articulate why a response receives a specific score to the exclusion of adjacent grades. Extensive experiments across datasets in physics, chemistry, and pedagogical content knowledge demonstrate that GUIDE significantly outperforms standard retrieval baselines. By focusing the model's attention on the precise edges of rubric, our approach shows exceptionally robust gains on borderline cases and improved rubric adherence. GUIDE paves the way for trusted, scalable assessment systems that align closely with human pedagogical standards.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。