arXiv:2605.04131physics.ed-phcs.AI2026-05被引 1

用对话框架解决AI在图文并茂的理科题中出错的问题。

A Dialogue-Based Framework for Correcting Multimodal Errors in AI-Assisted STEM Education

  • 设计对话流程引导AI逐步解析图文混合题目
  • 对话干预使错误率降低82%,图像理解错误全消除
  • 无需重训练,教师学生可立即使用

大型语言模型(LLMs)正推动个性化辅导普及,但其处理多模态内容的能力受限,影响了在科学、技术、工程和数学(STEM)教育中提供公平高质量支持的潜力。本研究评估了三个公开可用的LLM(Claude、Gemini和ChatGPT)在来自OpenStax数据库的多模态物理题上的表现,并与纯文本题结果对比。通过试点测试构建了基于实证的错误分类体系,随后评估了一种结构化多模态对话干预。所有模型在纯文本问题上准确率达96%以上;而在多模态问题上表现显著下降,出现所谓“多模态干扰效应”。错误分析识别出四类失败模式:视觉处理错误、上下文误读、数学计算错误和混合错误,其中视觉处理错误最常见。结构化对话干预总体纠正了82%的错误,对所有模型的视觉处理错误纠正率达100%。该方法无需模型重训练,教育者和学生可直接应用,提升图像丰富型STEM内容下AI辅导的可靠性,促进高质量学习支持的公平获取。

原文摘要 · Abstract (English)

Large Language Models (LLMs) are democratizing access to personalized tutoring; however, their effectiveness is hindered by challenges in processing multimodal content, which limits AI's potential to provide equitable, high-quality STEM support. This study evaluates LLM performance on multimodal physics problems, identifies specific failure modes through an empirical error taxonomy, and tests practical interventions designed to overcome multimodal processing limitations. We assessed three publicly available LLMs (Claude, Gemini, and ChatGPT) on multimodal physics problems from the OpenStax database and compared the results with text-only performance. An empirically derived error taxonomy was developed through pilot testing, followed by evaluation of a structured multimodal dialogue intervention. All three models achieved near-ceiling accuracy (96%) on text-only physics problems. Performance declined substantially on multimodal problems, consistent with what we term the Multimodal Interference Effect. Error analysis identified four failure modes: visual processing errors, context misinterpretation, mathematical computational errors, and hybrid errors, with visual processing errors being the most prevalent. The structured dialogue intervention corrected 82% of errors overall; visual processing errors were corrected at 100% across all models. Educators and students can implement these interventions immediately, requiring no model retraining, to improve AI tutoring reliability on image-rich STEM content, advancing equitable access to high-quality learning support.

AI教育多模态对话系统纠错

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。