用多智能体协作提升多模态作文评分的准确性与可信度
CAFES: A Collaborative Multi-Agent Framework for Multi-Granular Multimodal Essay Scoring
- 设计三个分工智能体:初评、反馈聚合、反思修正
- 相比人工标注,评分一致性提升21%(QWK)
- 适合需要高可信度评分的教育评估场景
自动作文评分(AES)在现代教育中至关重要,尤其在多模态评估日益普遍的背景下。传统方法在泛化能力和多模态感知上存在不足,而现有基于多模态大模型(MLLM)的方法常产生幻觉性理由和与人类判断不符的分数。为此,我们提出CAFES,首个专为AES设计的协作式多智能体框架。该框架包含三个专用智能体:初始评分器负责快速进行特质化评估;反馈池管理器聚合详细且有证据支持的优点;反思评分器则基于反馈迭代优化分数,以增强与人类判断的一致性。在使用前沿MLLM的大量实验中,相较于人工标注基准,平均相对提升达21%(QWK),尤其在语法和词汇多样性方面表现显著。CAFES为构建智能化多模态作文评分系统开辟了新路径。代码将在论文接受后公开。
原文摘要 · Abstract (English)
Automated Essay Scoring (AES) is crucial for modern education, particularly with the increasing prevalence of multimodal assessments. However, traditional AES methods struggle with evaluation generalizability and multimodal perception, while even recent Multimodal Large Language Model (MLLM)-based approaches can produce hallucinated justifications and scores misaligned with human judgment. To address the limitations, we introduce CAFES, the first collaborative multi-agent framework specifically designed for AES. It orchestrates three specialized agents: an Initial Scorer for rapid, trait-specific evaluations; a Feedback Pool Manager to aggregate detailed, evidence-grounded strengths; and a Reflective Scorer that iteratively refines scores based on this feedback to enhance human alignment. Extensive experiments, using state-of-the-art MLLMs, achieve an average relative improvement of 21% in Quadratic Weighted Kappa (QWK) against ground truth, especially for grammatical and lexical diversity. Our proposed CAFES framework paves the way for an intelligent multimodal AES system. The code will be available upon acceptance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。