通过合成数据微调,提升模型对双语写作的评分能力。
Improving Bilingual Capabilities of Language Models to Support Diverse Linguistic Practices in Education
- 用英语、西班牙语和西葡混用语合成数据微调模型。
- 微调后模型在三种语言上的评分表现显著提升。
- 适合关注双语教育公平与多语言模型设计的研究者。
大型语言模型(LLMs)在生成教育内容、提供教师反馈和减轻评估工作量方面具有潜力。然而,现有研究多聚焦于单语场景,较少探讨其在双语环境中的表现。本文研究了多语言大模型(MLLMs)在单语(仅英语、仅西班牙语)和双语(西葡混用语)学生写作中的有效性。我们构建了一个学习分析应用场景,评估模型对科学与社会科学概念解释的可接受性判断。结果发现,预训练模型在双语写作评分中存在显著偏差,表现劣于单语写作。随后,我们使用英语、西班牙语和西葡混用语生成的合成数据,对Llama 3.1和Mistral NeMo等开源模型进行微调。实验表明,微调后模型在三类语言上的表现均有显著提升。本研究凸显了通过数据增强提升多语言模型在双语学习者中有效性的潜力,并强调在教育领域设计语言模型时纳入非英语语言的重要性。
原文摘要 · Abstract (English)
Large language models (LLMs) offer promise in generating educational content, providing instructor feedback, and reducing teacher workload on assessments. While prior studies have focused on studying LLM-powered learning analytics, limited research has examined how effective LLMs are in a bilingual context. In this paper, we study the effectiveness of multilingual large language models (MLLMs) across monolingual (English-only, Spanish-only) and bilingual (Spanglish) student writing. We present a learning analytics use case that details LLM performance in assessing acceptable and unacceptable explanations of Science and Social Science concepts. Our findings reveal a significant bias in the grading performance of pre-trained models for bilingual writing compared to English-only and Spanish-only writing. Following this, we fine-tune open-source MLLMs including Llama 3.1 and Mistral NeMo using synthetic datasets generated in English, Spanish, and Spanglish. Our experiments indicate that the models perform significantly better for all three languages after fine-tuning with bilingual data. This study highlights the potential of enhancing MLLM effectiveness to support authentic language practices amongst bilingual learners. It also aims to illustrate the value of incorporating non-English languages into the design and implementation of language models in education.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。