arXiv:2509.14834cs.CL2025-09EMNLP被引 3

让AI模型像专家讨论一样评分作文,提升自动化评分准确性。

LLM Agents at the Roundtable: A Multi-Perspective and Dialectical Reasoning Framework for Essay Scoring

  • 构建多个角色化的AI评分员,分别从不同角度独立打分。
  • 通过模拟辩论式协商,综合多视角意见,得分更接近真人评判。
  • 在ASAP数据集上比传统方法提升34.86%的评分一致性。

大型语言模型(LLMs)为自动化作文评分(AES)带来了新范式,这一自然语言处理在教育中的长期应用仍面临实现人类级多视角理解与判断的挑战。本文提出Roundtable Essay Scoring(RES),一种基于LLM的多代理评估框架,在零样本设置下实现精准且符合人类标准的评分。RES构建多个基于不同提示和主题上下文的评价代理,每个代理独立生成基于特征的评分量表并执行多视角评估。随后,通过模拟圆桌会议式的讨论,利用辩证推理过程整合个体评价,生成最终的综合评分,更贴近人类评判。通过促进具有多样化评价视角的代理间协作与共识,RES优于以往零样本AES方法。在ASAP数据集上使用ChatGPT和Claude进行实验表明,RES相较直接提示法(Vanilla)在平均加权肯德尔协调系数(QWK)上最高提升34.86%。

原文摘要 · Abstract (English)

The emergence of large language models (LLMs) has brought a new paradigm to automated essay scoring (AES), a long-standing and practical application of natural language processing in education. However, achieving human-level multi-perspective understanding and judgment remains a challenge. In this work, we propose Roundtable Essay Scoring (RES), a multi-agent evaluation framework designed to perform precise and human-aligned scoring under a zero-shot setting. RES constructs evaluator agents based on LLMs, each tailored to a specific prompt and topic context. Each agent independently generates a trait-based rubric and conducts a multi-perspective evaluation. Then, by simulating a roundtable-style discussion, RES consolidates individual evaluations through a dialectical reasoning process to produce a final holistic score that more closely aligns with human evaluation. By enabling collaboration and consensus among agents with diverse evaluation perspectives, RES outperforms prior zero-shot AES approaches. Experiments on the ASAP dataset using ChatGPT and Claude show that RES achieves up to a 34.86% improvement in average QWK over straightforward prompting (Vanilla) methods.

作文评分多智能体辩证推理零样本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。