arXiv:2606.06754cs.MAcs.CL2026-06

用多智能体辩论+检索增强,实现无需训练的作文评分

MADRAG: Multi-Agent Debate with Retrieval-Augmented Generation for Training-Free Analytic Essay Scoring

论文配图:MADRAG: Multi-Agent Debate with Retrieval-Augmented Generation for Training-Free Analytic Essay Scoring
图 1 · 摘自论文原文
  • 三个智能体分工:辩护者找优点,质疑者挑毛病,裁判综合打分
  • 相比提示工程方法,评分更稳定;接近有监督模型效果
  • 检索优秀范例辅助裁判,提升评分准确性,适合教育评估场景

我们提出MADRAG,一种无需训练的分析性作文评分框架,结合多智能体推理与检索增强的基准对齐。不同于易受偏见影响的标准大模型评分方法,MADRAG将评估分解为互动过程:辩护者识别优点,质疑者批判缺点,裁判综合双方论点生成最终分数。关键在于,裁判通过检索与评分标准对齐的范例进行校准,实现基于样本比较的精准判断。实验表明,MADRAG显著优于提示工程基线,且在不需任务特定训练的情况下逼近有监督系统表现。消融研究显示,检索带来校准提升,辩论则增强对高阶特征的推理能力。结果凸显结构化交互与外部记忆在可靠大模型评估中的互补作用。

原文摘要 · Abstract (English)

We present MADRAG, a training-free framework for analytic essay scoring that combines multi-agent reasoning with retrieval-augmented grounding. Unlike standard LLM-as-judge approaches, which are prone to bias and unstable scoring, MADRAG decomposes evaluation into an interactive process: an Advocate identifies strengths, a Skeptic critiques weaknesses, and a Judge aggregates their arguments into a final score. Crucially, the Judge is augmented with rubric-aligned exemplar retrieval, enabling calibration through comparison with scored examples. Our results show that MADRAG significantly outperforms prompt-based baselines while approaching the performance of supervised systems without requiring task-specific training. Ablation studies demonstrate that retrieval drives calibration gains, while debate improves reasoning on higher-level traits. Our findings highlight the complementary roles of structured interaction and external memory in reliable LLM-based evaluation.

作文评分多智能体检索增强

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。