arXiv:2502.13789cs.CV2025-02被引 8

让AI不只判对错,还能分析学生错误原因并给出个性化建议

From Correctness to Comprehension: AI Agents for Personalized Error Diagnosis in Education

  • 构建多模态错误诊断数据集MathCCS,含真实题目与专家标注错误类型
  • 现有主流模型错误分类准确率不足30%,反馈质量平均低于4分(满分10分)
  • 提出多智能体协作框架,结合历史数据与实时分析提升诊断精准度

大型语言模型(如GPT-4)在数学推理任务(如GSM8K)上已接近完美表现,但在个性化教育中的应用受限于过度关注正确性而忽视错误诊断与反馈生成。为解决此问题,本文提出三项贡献:首先,构建多模态基准测试集MathCCS,包含真实问题、专家标注的错误类别及纵向学生数据;对Qwen2-VL、LLaVA-OV、Claude-3.5-Sonnet和GPT-4o等前沿模型的评估显示,其错误分类准确率均未超过30%,反馈质量平均低于4/10,显著低于人类水平。其次,提出基于历史数据的序列化错误分析框架,以追踪学习趋势并提升诊断精度。最后,设计多智能体协作框架,融合时间序列分析智能体与多模态大模型智能体,实现动态优化的错误分类与反馈生成。整体方案为个性化教育提供了可扩展的技术平台。

原文摘要 · Abstract (English)

Large Language Models (LLMs), such as GPT-4, have demonstrated impressive mathematical reasoning capabilities, achieving near-perfect performance on benchmarks like GSM8K. However, their application in personalized education remains limited due to an overemphasis on correctness over error diagnosis and feedback generation. Current models fail to provide meaningful insights into the causes of student mistakes, limiting their utility in educational contexts. To address these challenges, we present three key contributions. First, we introduce \textbf{MathCCS} (Mathematical Classification and Constructive Suggestions), a multi-modal benchmark designed for systematic error analysis and tailored feedback. MathCCS includes real-world problems, expert-annotated error categories, and longitudinal student data. Evaluations of state-of-the-art models, including \textit{Qwen2-VL}, \textit{LLaVA-OV}, \textit{Claude-3.5-Sonnet} and \textit{GPT-4o}, reveal that none achieved classification accuracy above 30\% or generated high-quality suggestions (average scores below 4/10), highlighting a significant gap from human-level performance. Second, we develop a sequential error analysis framework that leverages historical data to track trends and improve diagnostic precision. Finally, we propose a multi-agent collaborative framework that combines a Time Series Agent for historical analysis and an MLLM Agent for real-time refinement, enhancing error classification and feedback generation. Together, these contributions provide a robust platform for advancing personalized education, bridging the gap between current AI capabilities and the demands of real-world teaching.

教育AI错误诊断多智能体个性化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。