arXiv:2512.10785physics.ed-phcs.AI2025-12

基于证据中心设计的LLM反馈系统助力物理解题,但20%错误难被学生察觉。

Developing an LLM-Based Feedback System Grounded in Evidence-Centered Design to Support Physics Problem Solving

  • 以证据中心设计为框架构建LLM反馈系统,提升物理题解的个性化指导。
  • 德国物理奥赛测试显示反馈正确率高,但20%存在错误且学生难以识别。
  • 提醒警惕盲目依赖AI反馈,适合教育技术研究者与教学AI开发者参考。

生成式AI为个性化自适应学习带来新机遇,例如基于大语言模型(LLM)的反馈系统。尽管LLM在简单概念任务中可生成事实正确的反馈,但在需要高级领域知识的任务(如物理问题求解)中,提供高质量反馈仍具挑战。本研究提出一种基于证据中心设计的物理问题求解LLM反馈系统,并在德国物理奥赛中进行首次评估。参与者对每道题的反馈实用性与正确性进行评分,结果显示反馈普遍被认为有用且高度正确。然而深入分析发现,20%的反馈存在错误,且多数学生未能察觉。本文讨论了过度依赖LLM反馈的风险,并提出了未来生成更自适应、更可靠反馈的方向。

原文摘要 · Abstract (English)

Generative AI offers new opportunities for individualized and adaptive learning, e.g., through large language model (LLM)-based feedback systems. While LLMs can produce factually correct feedback for relatively straightforward conceptual tasks, delivering high-quality feedback for tasks that require advanced domain expertise, such as physics problem solving, remains a substantial challenge. This study presents the design and implementation of an LLM-based feedback system for physics problem solving grounded in evidence-centered design and reports a first evaluation within the German Physics Olympiad. Participants rated the usefulness and correctness of the generated feedback for each implemented problem. The collected ratings indicate that the feedback was generally perceived as useful and highly correct. However, an in-depth analysis revealed that the feedback contained errors in 20% of cases; errors that often went unnoticed by the students. We discuss the risks associated with uncritical reliance on LLM-based feedback and outline potential directions for generating more adaptive and reliable LLM-based feedback in the future.

AI教育物理学习LLM反馈

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。