arXiv:2605.29274cs.CL2026-05

让大模型自己学会评分规则,无需人工编写标准。

Learnable Assessment Skills for LLM-based Automated Scoring: Rubric Construction via Iterative Optimization

论文配图:Learnable Assessment Skills for LLM-based Automated Scoring: Rubric Construction via Iterative Optimization
图 1 · 摘自论文原文
  • 用迭代优化方法将评分规则拆解为可学习的通用流程
  • 在10个题项上表现优于人工制定的评分标准
  • 学成的技能能跨题目迁移,兼顾通用与特定模式

基于大模型的自动评分已接近人类水平,但扩展到新任务时仍受限于需人工配置评分标准等上游环节。人类专家通过长期实践形成高效评估策略。本文探讨大模型能否从评分经验中直接学习类似策略,并提出“评估技能”概念:一种与题目无关、以自然语言描述的流程化知识,指导模型完成评分流程中的特定阶段。以评分标准构建为首个实例,提出一个迭代框架,将技能分解为固定结构和可学习的通用规则,通过大模型诊断评分错误并经验证筛选不断优化规则。该框架无需专家编写评分标准。在所有10个ASAP-SAS题项上,优化后的技能显著提升评分性能,常超越数据集提供的专家标准。跨题项迁移实验表明,所学技能同时捕捉通用规律与题项特异性模式。

原文摘要 · Abstract (English)

LLM-based automated scoring approaches near-human performance, but scaling to new tasks remains bottlenecked by the per-item human configuration of upstream stages such as rubric construction. Human experts bypass this bottleneck through evaluation heuristics developed over extensive practice. We ask whether LLMs can learn similar heuristics directly from scoring experience, and formalize this as the concept of assessment skills: item-independent natural-language procedural knowledge that guides LLMs through specific stages of the scoring workflow. Focusing on rubric construction as a first instantiation, we propose an iterative framework that decomposes a skill into a fixed scaffold and learnable item-agnostic rules, refining the rules through LLM-driven diagnosis of scoring errors and validation-gated selection. The framework requires no expert-written rubric. On all ten ASAP-SAS items, optimized skills substantially improve LLM-based scoring and frequently surpass the dataset-provided expert rubric. Cross-item transfer experiments further reveal that learned skills capture both generalizable and item-specific patterns.

自动评分大模型评估技能迭代优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。