实测发现学生常跳过AI导师的引导步骤,暴露教学设计与真实学习目标的错配。
Rethinking Scaffolding in LLM Tutors: The Interactional Mismatch Between Benchmarks and Real-World Deployments

- 构建双指标评估体系,量化聊天机器人引导与学生实际响应程度。
- 9490条对话分析显示,真实场景中学生采纳引导的比例远低于基准测试假设。
- 提示未来评测应关注学生主导的互动模式,而非单向依赖引导机制。
AI导师评测中的核心教育价值之一是“支架式引导”:通过渐进步骤引导学生解题。然而,现有嵌入引导行为的方法隐含一个假设:学生会接受并参与该引导过程。为检验这一假设,我们提出包含两个指标的评估流程——聊天机器人引导度与学生采纳度,并在涵盖九个数据集、共9490条对话的AI导师评测与真实教育聊天机器人部署中进行应用。分析发现,尽管基准测试假设高引导、高采纳环境,但真实场景中学生整体采纳率较低,频繁跳过机器人的教学框架,转而按自身学习目标推进对话,且人际成本极低。我们认为跳过引导未必有害,反而常反映出聊天机器人教学框架与学生学习目标之间的不匹配。为有效评估聊天机器人辅助效果,未来的评测必须突破‘学生必然采纳引导’的预设,转而考察其在多样学习情境和学生主导交互模式下的适应能力。
原文摘要 · Abstract (English)
A central pedagogical value evaluated in AI tutor benchmarks is scaffolding: guiding students through graduated steps toward a solution. Alignment and evaluation methods for embedding scaffolding behaviour into chatbots, however, rest on an implicit assumption: that students will take up the scaffolding and engage in the conversation. To examine whether this assumption holds, we introduce an evaluation pipeline around two metrics - Chatbot Scaffolding and Student Uptake - and apply them across nine datasets of 9,490 chats, spanning AI tutor benchmarks and real-world deployments of educational chatbots. Our analysis reveals that while benchmarks assume a high-scaffolding, high-student-uptake environment, students in real-world settings exhibit lower levels of uptake overall - frequently bypassing the chatbot's pedagogical framing to drive the interaction toward their own learning goals at little interpersonal cost. We argue that bypassing scaffolding is not necessarily detrimental; rather, it frequently highlights a mismatch between a chatbot's pedagogical framing and the student's learning goals. To meaningfully evaluate the effectiveness of a chatbot's assistance, future benchmarks must move beyond the assumption that students will simply take up the scaffolding, and instead evaluate how these chatbots navigate diverse learning contexts and student-driven interaction patterns.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。