多智能体系统更擅长识别差作文,单模型更适合一般评分。
Specialists or Generalists? Multi-Agent and Single-Agent LLMs for Essay Grading
- 用三个专业代理+总指挥分工评分,按评分标准协同决策。
- 少样本微调仅需每档2例,就能让评分一致性提升26%。
- 多智能体适合筛查学困生,单模型则适合低成本普适评估。
自动作文评分系统越来越多依赖大语言模型,但架构选择如何影响不同质量作文的评分表现仍不明确。本研究基于ASAP 2.0语料库,对比了单智能体与多智能体架构在作文评分中的表现。多智能体系统将评分分解为内容、结构、语言三个专业代理,由总指挥代理依据评分标准(含否决权和分数上限)协调。在零样本和少样本条件下使用GPT-5.1测试。结果表明,多智能体系统在识别弱作文方面显著更优,而单智能体在中等质量作文上表现更好。两者对高分作文均表现不佳。关键发现是:少样本校准是性能主导因素——每档仅提供两个示例,即可使整体评分一致性(QWK)提升约26%。这提示架构选择应匹配实际需求:多智能体适合诊断性筛查学业风险学生,单智能体则为通用评估提供高性价比方案。
原文摘要 · Abstract (English)
Automated essay scoring (AES) systems increasingly rely on large language models, yet little is known about how architectural choices shape their performance across different essay quality levels. This paper evaluates single-agent and multi-agent LLM architectures for essay grading using the ASAP 2.0 corpus. Our multi-agent system decomposes grading into three specialist agents (Content, Structure, Language) coordinated by a Chairman Agent that implements rubric-aligned logic including veto rules and score capping. We test both architectures in zero-shot and few-shot conditions using GPT-5.1. Results show that the multi-agent system is significantly better at identifying weak essays while the single-agent system performs better on mid-range essays. Both architectures struggle with high-quality essays. Critically, few-shot calibration emerges as the dominant factor in system performance -- providing just two examples per score level improves QWK by approximately 26% for both architectures. These findings suggest architectural choice should align with specific deployment priorities, with multi-agent AI particularly suited for diagnostic screening of at-risk students, while single-agent models provide a cost-effective solution for general assessment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。