多智能体系统提升编程作业评分准确性和一致性
AGACCI : Affiliated Grading Agents for Criteria-Centric Interface in Educational Coding Contexts
- 将评估任务拆分给多个协作智能体,各司其职
- 在360份研究生代码作业上表现优于单模型基线
- 适合需要结构化评判的教育评测场景
近期人工智能辅助教育的发展推动了视觉语言模型(VLM)在学术评估中的应用,尤其适用于需定量与定性双重评价的任务。然而,现有VLM方法在处理包含可执行代码和可测量输出的复杂教育成果时仍存在挑战,难以实现结构化推理与评价标准对齐。本文提出AGACCI,一种多智能体系统,通过分布式专业评估角色提升代码类作业评估的准确性、可解释性与一致性。为验证框架,我们收集了60名参与者提交的360份研究生级别代码作业,每份均由领域专家标注二元评分表与定性反馈。实验表明,AGACCI在评分准确性、反馈相关性、一致性与连贯性方面均优于单一GPT基线,同时保持了专家评估的教学意图与评价深度。尽管不同任务类型表现有所差异,但结果凸显多智能体系统在可扩展、情境感知教育评估中的潜力。
原文摘要 · Abstract (English)
Recent advances in AI-assisted education have encouraged the integration of vision-language models (VLMs) into academic assessment, particularly for tasks that require both quantitative and qualitative evaluation. However, existing VLM based approaches struggle with complex educational artifacts, such as programming tasks with executable components and measurable outputs, that require structured reasoning and alignment with clearly defined evaluation criteria. We introduce AGACCI, a multi-agent system that distributes specialized evaluation roles across collaborative agents to improve accuracy, interpretability, and consistency in code-oriented assessment. To evaluate the framework, we collected 360 graduate-level code-based assignments from 60 participants, each annotated by domain experts with binary rubric scores and qualitative feedback. Experimental results demonstrate that AGACCI outperforms a single GPT-based baseline in terms of rubric and feedback accuracy, relevance, consistency, and coherence, while preserving the instructional intent and evaluative depth of expert assessments. Although performance varies across task types, AGACCI highlights the potential of multi-agent systems for scalable and context-aware educational evaluation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。