arXiv:2602.04604cs.CL2026-02被引 1

提出基于评分维度的作文质量评估方法,提升教育反馈可解释性。

Beyond Holistic Scores: Automatic Trait-Based Quality Scoring of Argumentative Essays

  • 用小模型+上下文学习和大模型+有序回归,实现多维度作文评分。
  • 显式建模分数有序性,与人工评分一致性显著提升。
  • 开源小模型无需微调即表现良好,适合隐私敏感场景。

自动化作文评分系统传统上依赖整体评分,限制了其在复杂文体如议论文中的教学价值。教师与学生需要符合教学目标和评分量表的可解释性、维度级反馈。本文研究基于特质的自动议论文评分,采用两种互补建模范式:(1) 使用小型开源大模型进行结构化上下文学习;(2) 基于监督编码器的BigBird模型,结合CORAL风格的有序回归,优化长文本理解。我们在ASAP++数据集上进行系统评估,该数据集涵盖五个质量维度的作文评分,全面覆盖核心论证维度。大模型通过秩一致的CORAL框架显式建模分数序数关系,小模型则使用设计好的、对齐评分量表的上下文示例进行提示,并请求反馈与置信度。结果表明,显式建模分数有序性显著提升与人工评分的一致性,优于大模型、名义分类与回归基线。这一发现强调模型目标需与评分量表语义对齐的重要性。同时,小型开源大模型在无需任务微调的情况下取得有竞争力的表现,尤其在推理相关维度上表现突出,且支持透明、隐私保护与本地部署的评估场景。研究为构建可解释、量表对齐的AI教育系统提供了方法论、建模与实践启示。

原文摘要 · Abstract (English)

Automated Essay Scoring systems have traditionally focused on holistic scores, limiting their pedagogical usefulness, especially in the case of complex essay genres such as argumentative writing. In educational contexts, teachers and learners require interpretable, trait-level feedback that aligns with instructional goals and established rubrics. In this paper, we study trait-based Automatic Argumentative Essay Scoring using two complementary modeling paradigms designed for realistic educational deployment: (1) structured in-context learning with small open-source LLMs, and (2) a supervised, encoder-based BigBird model with a CORAL-style ordinal regression formulation, optimized for long-sequence understanding. We conduct a systematic evaluation on the ASAP++ dataset, which includes essay scores across five quality traits, offering strong coverage of core argumentation dimensions. LLMs are prompted with designed, rubric-aligned in-context examples, along with feedback and confidence requests, while we explicitly model ordinality in scores with the BigBird model via the rank-consistent CORAL framework. Our results show that explicitly modeling score ordinality substantially improves agreement with human raters across all traits, outperforming LLMs and nominal classification and regression-based baselines. This finding reinforces the importance of aligning model objectives with rubric semantics for educational assessment. At the same time, small open-source LLMs achieve a competitive performance without task-specific fine-tuning, particularly for reasoning-oriented traits, while enabling transparent, privacy-preserving, and locally deployable assessment scenarios. Our findings provide methodological, modeling, and practical insights for the design of AI-based educational systems that aim to deliver interpretable, rubric-aligned feedback for argumentative writing.

作文评分大模型教育AI有序回归

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。