测试AI助教何时干预最有效,发现其常过早给答案,反而不利于学习。
AI Assistants Overassist

- 设计模拟学生解题场景,评估AI教师干预时机与频率
- AI比人类更频繁、更早干预,且多给完整答案而非提示
- 当前AI倾向短期成功,忽视长期思维能力培养
大型语言模型(LLMs)越来越多地被用作导师和思维伙伴,帮助用户解决问题。尽管AI的指导能促进思考与学习,但效果取决于干预方式——过早或过频的介入可能阻碍真正的学习与认知参与。然而,目前对AI在解题过程中如何做出干预决策仍缺乏理解。本文提出Int-Bench,一个基于仿真的基准测试,用于评估LLM在学习过程中的干预行为。该基准模拟一名“学生”解题,由“教师”监控其推理并决定是否、何时以及如何干预。我们在代码调试、数学问题和脑筋急转弯三个领域评估了LLM教师在干预频率与时机上的表现,及其对即时任务成功率和泛化能力的影响。结果表明,与人类相比,LLM更频繁、更早干预,且倾向于提供完整解决方案而非针对性提示。这些发现表明,当前的LLM助手往往优化于短期成功,而非支持深度学习所需的推理过程。
原文摘要 · Abstract (English)
Large language models (LLMs) are increasingly used as tutors and thought partners, helping users reason through problems. While guidance from AI assistants can scaffold thinking and foster learning, such benefits depend on how they help--for instance, intervening too early or too frequently may hinder true learning and cognitive engagement. Yet how AI systems navigate intervention decisions during problem-solving remains poorly understood. Here, we introduce Int-Bench, a simulation-based benchmark for evaluating LLM interventions during learning. Int-Bench simulates a "student" solving a problem while a "teacher" monitors the student's reasoning and decides whether, when, and how to intervene. Across three domains--code debugging, mathematics, and brain teasers--we evaluate LLM teachers on the frequency and timing of interventions, as well as their impact on both immediate task success and generalization to new problems. We also compare LLMs to humans, finding that LLMs intervene more frequently and earlier than humans. Moreover, in contrast to humans, they tend to provide complete solutions rather than targeted hints. These findings suggest that current LLM assistants often optimize for short-term success rather than supporting the reasoning processes needed for deeper learning and long-term success.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。