研究大模型在西语教学中如何因提示导致能力错配
Alignment Drift in CEFR-prompted LLMs for Interactive Spanish Tutoring
- 用CEFR分级提示控制模型输出难度,模拟师生对话
- 7B到12B参数模型在三水平对话中出现持续能力偏移
- 揭示提示依赖性弱点,适合教育AI评估研究者参考
本文探究大语言模型(LLMs)作为第二语言学习自适应导师的潜力。特别考察系统提示是否能可靠约束模型仅生成符合学生能力水平的内容。我们使用指令微调的开源LLM(参数量7B至12B)模拟完整的西班牙语师生对话,让模型在导师与学生角色间交替,各自维护独立聊天历史。通过导师模型输出,评估基于CEFR等级(A1、B1、C1)的提示对文本难度控制的效果。结果表明,尽管系统提示可短期约束输出,但提示本身过于脆弱,在长期交互中会引发‘对齐漂移’现象。研究为大模型在个性化、能力匹配的自适应教学中的可行性提供见解,并提出一种无需真人参与的低成本模型性能评估方法。
原文摘要 · Abstract (English)
This paper investigates the potentials of Large Language Models (LLMs) as adaptive tutors in the context of second-language learning. In particular, we evaluate whether system prompting can reliably constrain LLMs to generate only text appropriate to the student's competence level. We simulate full teacher-student dialogues in Spanish using instruction-tuned, open-source LLMs ranging in size from 7B to 12B parameters. Dialogues are generated by having an LLM alternate between tutor and student roles with separate chat histories. The output from the tutor model is then used to evaluate the effectiveness of CEFR-based prompting to control text difficulty across three proficiency levels (A1, B1, C1). Our findings suggest that while system prompting can be used to constrain model outputs, prompting alone is too brittle for sustained, long-term interactional contexts - a phenomenon we term alignment drift. Our results provide insights into the feasibility of LLMs for personalized, proficiency-aligned adaptive tutors and provide a scalable method for low-cost evaluation of model performance without human participants.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。