让大模型对话更适配初学者,自动调节语言难度
Toward Beginner-Friendly LLMs for Language Learning: Controlling Difficulty in Conversation
- 用可控生成技术降低大模型输出的语言复杂度
- 初学者理解率从39.4%提升至83.3%
- 新指标TMR能准确衡量难懂词占比,适合研究者使用
与传统面对面语言学习相比,使用大语言模型(LLMs)进行对话练习提供了有前景的替代方案。然而,大多数LLM生成的内容接近母语水平,难以满足初级学习者(CEFR A1-A2)的需求。本文探讨了可控生成技术是否可使LLM输出更适合初学者。通过自动评估指标和面向日本语大学生的学习者用户研究,我们发现仅靠提示(prompting)无效,但可控生成技术能显著提升输出可理解性(从39.4%提升至83.3%)。我们还提出了一种新的逐标记评估指标——令牌错失率(Token Miss Rate, TMR),用于量化每句话中不可理解的标记比例,其结果与人工判断高度相关。为支持未来在人工智能辅助语言学习方面的研究,我们公开了代码、模型、标注工具及数据集。
原文摘要 · Abstract (English)
Practicing conversations with large language models (LLMs) presents a promising alternative to traditional in-person language learning. However, most LLMs generate text at a near-native level of complexity, making them ill-suited for first and second-year beginner learners (CEFR: A1-A2). In this paper, we investigate whether controllable generation techniques can adapt LLM outputs to better support beginners. We evaluate these methods through both automatic metrics and a user study with university-level learners of Japanese. Our findings show that while prompting alone fails, controllable generation techniques can successfully improve output comprehensibility for beginner speakers (from 39.4% to 83.3%). We further introduce a new token-level evaluation metric, Token Miss Rate (TMR), that quantifies the proportion of incomprehensible tokens per utterance and correlates strongly with human judgments. To support future research in AI-assisted language learning, we release our code, models, annotation tools, and dataset.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。