用对话数据训练可控制大模型教学风格的向量,让AI导师更像真人。
Letting Tutor Personas Speak Up for LLMs: Learning Steering Vectors from Dialogue via Preference Optimization
- 通过偏好优化从真实师生对话中学习教学风格向量
- 使AI回应更贴近真人导师,提升评分与语义匹配度
- 可解释性强,适合个性化教育、智能辅导系统研究
随着大语言模型(LLMs)在生成式人工智能中的兴起,其在教学场景中的应用日益突出。以往基于LLM的教学研究通常只学习单一教学策略,未能体现教学风格的多样性。现实中,教师会根据学生需求动态调整支架程度、指导性、反馈方式和情感支持,这些差异显著影响对话互动与学生参与度。本文探索如何利用人类师生对话中嵌入的教师人格特征来引导LLM行为,而无需显式提示。我们采用偏好优化训练一个激活空间方向(即导向向量),使模型输出趋近特定教师人格。实验表明,该向量能有效捕捉不同情境下的教学风格差异,在语义上更接近真实教师话语,并在偏好评估中表现更优,同时保持较高的词汇相似性。对学习到的缩放系数分析揭示了可解释的教师行为模式结构,证明激活空间引导是一种有效且可解释的、基于人类对话数据控制教学风格的方法。
原文摘要 · Abstract (English)
With the emergence of large language models (LLMs) as a powerful class of generative artificial intelligence (AI), their use in tutoring has become increasingly prominent. Prior works on LLM-based tutoring typically learn a single tutor policy and do not capture the diversity of tutoring styles. In real-world tutor-student interactions, pedagogical intent is realized through adaptive instructional strategies, with tutors varying the level of scaffolding, instructional directiveness, feedback, and affective support in response to learners' needs. These differences can all impact dialogue dynamics and student engagement. In this paper, we explore how tutor personas embedded in human tutor-student dialogues can be used to guide LLM behavior without relying on explicitly prompted instructions. We train a steering vector using preference optimization: an activation-space direction that guides model responses toward specific tutor personas. We find that this steering vector captures tutor-specific variation across dialogue contexts, improving semantic alignment with ground-truth tutor utterances and increasing preference-based evaluations, while largely preserving lexical similarity. Analysis of the learned scaling coefficients further reveals interpretable structure across tutors, corresponding to consistent differences in tutoring behavior. These results demonstrate that activation steering offers an effective and interpretable way for controlling tutor-specific variation in LLMs using signals derived directly from human dialogue data.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。