用多智能体系统让AI反馈可互动,支持学生追问与自我修正。
REFINE: Real-world Exploration of Interactive Feedback and Student Behaviour
- 基于小模型构建交互式反馈系统,支持追问与上下文响应。
- 自动评估显示,人工对齐裁判引导重生成显著提升反馈质量。
- 真实课堂实测表明系统能引导学生后续提问方向,适合教育科技研发者。
形成性反馈是有效学习的核心,但大规模提供及时、个性化的反馈仍具挑战。尽管近期研究探索用大语言模型(LLMs)自动化反馈,现有系统大多将反馈视为静态、单向的产物,难以支持解释、澄清或后续互动。本文提出REFINE,一个基于小型开源LLM、可本地部署的多智能体反馈系统,将反馈视为交互过程。REFINE结合了教学法驱动的反馈生成代理、基于人类对齐裁判引导的重生成循环,以及支持上下文感知、可操作响应的学生追问自省工具调用代理。通过受控实验和在本科生计算机科学课程中的真实课堂部署进行评估。自动评估显示,裁判引导的重生成显著提升反馈质量;交互代理生成的响应在效率与质量上可媲美先进闭源模型。真实学生交互分析揭示了不同的参与模式,并表明系统生成的反馈会系统性引导后续学生提问。研究结果证明了多智能体、工具增强型反馈系统在实现可扩展交互反馈方面的可行性与有效性。
原文摘要 · Abstract (English)
Formative feedback is central to effective learning, yet providing timely, individualised feedback at scale remains a persistent challenge. While recent work has explored the use of large language models (LLMs) to automate feedback, most existing systems still conceptualise feedback as a static, one-way artifact, offering limited support for interpretation, clarification, or follow-up. In this work, we introduce REFINE, a locally deployable, multi-agent feedback system built on small, open-source LLMs that treats feedback as an interactive process. REFINE combines a pedagogically-grounded feedback generation agent with an LLM-as-a-judge-guided regeneration loop using a human-aligned judge, and a self-reflective tool-calling interactive agent that supports student follow-up questions with context-aware, actionable responses. We evaluate REFINE through controlled experiments and an authentic classroom deployment in an undergraduate computer science course. Automatic evaluations show that judge-guided regeneration significantly improves feedback quality, and that the interactive agent produces efficient, high-quality responses comparable to a state-of-the-art closed-source model. Analysis of real student interactions further reveals distinct engagement patterns and indicates that system-generated feedback systematically steers subsequent student inquiry. Our findings demonstrate the feasibility and effectiveness of multi-agent, tool-augmented feedback systems for scalable, interactive feedback.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。