LLM导师难以区分正确、有效但次优和错误解法,影响精准辅导。
Confirming Correct, Missing the Rest: LLM Tutoring Agents Struggle Where Feedback Matters Most

- 用知识图谱生成真实答案,测试7个LLM在命题逻辑中的反馈能力。
- 对正确步骤识别率高,但误判有效但低效解法为错误,错认错误解法为正确。
- 诊断不准导致教学建议无效,适合用于对话支持而非核心判断。
高效辅导需区分最优解、有效但次优解与错误解,这是智能辅导系统的核心,但尚未在基于大模型的导师中检验。随着大模型被探索为智能辅导系统的对话补充,评估其诊断精度至关重要。我们构建了一个包含10,836个解法-反馈对的基准,基于知识图谱生成的真实答案,在三种反馈条件下测试七个LLM反馈代理在命题逻辑中的表现。模型在最优步骤上接近完美,但系统性地过度拒绝有效但次优的推理,且过度认可错误解法,而这正是自适应辅导最关键的环节。这种失败在不同解法上下文中持续存在,表明是架构局限而非信息不足所致。此外,准确诊断并未可靠转化为可操作的教学反馈,暴露出诊断判断与教学效果之间的鸿沟。研究建议,应采用混合架构:由知识图谱驱动的模型负责诊断,而大模型则承担开放式的支架与对话支持。
原文摘要 · Abstract (English)
Effective tutoring requires distinguishing optimal, valid but suboptimal, and incorrect student solutions, a distinction central to intelligent tutoring systems (ITS) but untested for LLM-based tutors. As LLMs are increasingly explored as conversational complements to ITS, evaluating their diagnostic precision is essential. We present a benchmark of seven LLM feedback agents in propositional logic using knowledge-graph-derived ground truth across 10,836 solution--feedback pairs and three feedback conditions. Models achieved near-ceiling performance on optimal steps but systematically over-rejected valid but suboptimal reasoning and over-validated incorrect solutions, precisely where adaptive tutoring matters most. These failures persisted across models regardless of solution context, suggesting architectural rather than informational limits. Moreover, accurate diagnosis did not reliably produce pedagogically actionable feedback, revealing a gap between diagnostic judgment and instructional effectiveness. Our findings suggest that LLMs are better suited for hybrid architectures where KG-grounded models handle diagnosis while LLMs support open-ended scaffolding and dialogue.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。