arXiv:2511.11772cs.CYcs.AI2025-11AAAI

用五个角色智能体实现大规模教育反馈的公平与高效

Scaling Equitable Reflection Assessment in Education via Large Language Models and Role-Based Feedback Agents

  • 设计五类角色智能体协同评分与生成反馈
  • 评分接近专家水平,评论获人工评价为有帮助且共情
  • 适合大规模课程或资源匮乏环境中的教学支持

形成性反馈被广泛认为是促进学生学习最有效的手段之一,但在大规模或资源有限的课程中难以公平实施。教师常因时间、人力和精力不足,无法回复每位学生的反思,导致最需要支持的学生反而得不到及时反馈。本文提出一个基于理论的多智能体系统,由五类角色化大语言模型代理(评估者、公平监控者、元认知教练、聚合器、反思评审者)协同工作:先按统一量规打分,检查潜在偏见语言,加入引导自我反思的提示,并生成不超过120字的简洁反馈。系统包含简单的公平性检测机制,通过比较高低分组的评分误差,帮助教师监控并控制准确性差异。在为期12周的成人人工智能素养课程中评估显示,该系统评分接近专家一致性,训练过的评分员评价其生成的反馈具有帮助性、同理心且符合教学目标。结果表明,多智能体大模型系统可在人力无法企及的速度与规模下,实现高质量、公平的形成性反馈。更广泛而言,这项工作指向一个未来:无论课程规模或背景如何,都能实现丰富反馈的学习体验,推动教育公平、可及性与教学能力的长期目标。

原文摘要 · Abstract (English)

Formative feedback is widely recognized as one of the most effective drivers of student learning, yet it remains difficult to implement equitably at scale. In large or low-resource courses, instructors often lack the time, staffing, and bandwidth required to review and respond to every student reflection, creating gaps in support precisely where learners would benefit most. This paper presents a theory-grounded system that uses five coordinated role-based LLM agents (Evaluator, Equity Monitor, Metacognitive Coach, Aggregator, and Reflexion Reviewer) to score learner reflections with a shared rubric and to generate short, bias-aware, learner-facing comments. The agents first produce structured rubric scores, then check for potentially biased or exclusionary language, add metacognitive prompts that invite students to think about their own thinking, and finally compose a concise feedback message of at most 120 words. The system includes simple fairness checks that compare scoring error across lower and higher scoring learners, enabling instructors to monitor and bound disparities in accuracy. We evaluate the pipeline in a 12-session AI literacy program with adult learners. In this setting, the system produces rubric scores that approach expert-level agreement, and trained graders rate the AI-generated comments as helpful, empathetic, and well aligned with instructional goals. Taken together, these results show that multi-agent LLM systems can deliver equitable, high-quality formative feedback at a scale and speed that would be impossible for human graders alone. More broadly, the work points toward a future where feedback-rich learning becomes feasible for any course size or context, advancing long-standing goals of equity, access, and instructional capacity in education.

教育AI多智能体公平反馈

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。