MAGIC用五个智能体评估作文,提升高校写作评分与反馈质量。
MAGIC: Multi-Agent Argumentation and Grammar Integrated Critiquer
- 五类专用智能体分别评估论点、说服力、结构、词汇和语法。
- 在GRE作文数据集上评分与人工高度一致,接近完美。
- 适合需要高质量写作反馈的高等教育场景。
自动作文评分(AES)与自动作文反馈(AEF)系统旨在减轻教育评估中人工评分负担。然而,现有系统多侧重分数准确性,忽视反馈质量,且主要针对中小学水平写作进行评估。本文提出多智能体论辩与语法集成评阅框架MAGIC,通过五个专精智能体,分别评估作文对题目的契合度、说服力、组织结构、词汇使用和语法正确性,实现整体评分与详细反馈生成。为支持大学水平评估,我们收集了带有专家评分与反馈的研究生入学考试(GRE)练习作文数据集。MAGIC在该数据集上与人类评分者达成显著至近乎完美的评分一致性,优于基线大模型,且通过多智能体机制提升了可解释性。我们还将MAGIC的反馈生成能力与真实人工反馈及基线模型对比,结果表明其反馈质量与自然度均表现优异。
原文摘要 · Abstract (English)
Automated Essay Scoring (AES) and Automatic Essay Feedback (AEF) systems aim to reduce the workload of human raters in educational assessment. However, most existing systems prioritize numerical scoring accuracy over feedback quality and are primarily evaluated on pre-secondary school level writing. This paper presents Multi-Agent Argumentation and Grammar Integrated Critiquer (MAGIC), a framework using five specialized agents to evaluate prompt adherence, persuasiveness, organization, vocabulary, and grammar for both holistic scoring and detailed feedback generation. To support evaluation at the college level, we collated a dataset of Graduate Record Examination (GRE) practice essays with expert-evaluated scores and feedback. MAGIC achieves substantial to near-perfect scoring agreement with humans on the GRE data, outperforming baseline LLM models while providing enhanced interpretability through its multi-agent approach. We also compare MAGIC's feedback generation capabilities against ground truth human feedback and baseline models, finding that MAGIC achieves strong feedback quality and naturalness.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。