arXiv:2608.07509cs.HCcs.AI2026-08

让大模型导师按需调整教学策略,实时控制辅导行为。

PIVOT: Preference-based Intervention Vectors for Pedagogical Tutor Steering

论文配图:PIVOT: Preference-based Intervention Vectors for Pedagogical Tutor Steering
图 1 · 摘自论文原文
  • 通过偏好学习生成干预向量,动态调节模型教学动作。
  • 在30位教师测试中,73.3%更偏好可调控的对话。
  • 支持策略组合与迁移,适合教育类AI系统开发。

大语言模型在对话式辅导中应用日益广泛,但有效辅导不仅需要正确答案,还需适时提供引导、提示、反馈、解释或引发反思。现有提示和训练方法虽提升教学对齐度,却缺乏推理时对教学策略的可靠控制。本文提出PIVOT,一种针对冻结大模型导师的激活调控框架,通过在线学习偏好干预向量实现教学策略控制。PIVOT采用七类教学动作分类体系,结合生成-标注-优化循环,由人工验证的LLM评判器识别目标动作与混淆非目标动作,构建偏好对以实现多层残差流调控。在保留相关性与流畅性的前提下,该方法在保留数据及跨领域数据上均有效控制教学动作,且调控方向可在推理时缩放、迁移与组合。30位教师参与的用户研究显示,73.3%参与者更偏好经调控的对话,认为控制方式清晰、可用且具有教学意义。

原文摘要 · Abstract (English)

LLMs are increasingly used for conversational tutoring, but effective tutoring requires more than correct answers. Tutors must choose when to scaffold reasoning, hint, give feedback, explain, or invite reflection. Existing prompting and training methods improve pedagogical alignment, but lack reliable inference-time control over pedagogical strategies. We introduce PIVOT, an activation-steering framework that learns preference-based intervention vectors online for frozen LLM tutors. PIVOT uses a seven-category tutor-move taxonomy and a generate-label-optimise loop, where a human-validated LLM judge identifies target and confusable non-target moves to construct preference pairs for multi-layer residual-stream steering. Across held-out and out-of-domain tutoring data, PIVOT controls tutor moves while preserving relevance and fluency, and its directions can be scaled, transferred, and composed at inference time. In a user study with 30 teachers, 73.3% of participants preferred steered conversations over neutral baseline interactions using the same prompt, and rated the controls as clear, usable, and pedagogically meaningful.

教育AI提示工程策略控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。