选对多语种教师模型,能显著提升小模型训练效果
Polyglot Teachers: Evaluating Language Models for Multilingual Synthetic Data Generation
- 用数据质量+学生表现综合评估教师模型效能
- 小模型反而常比大模型更适合作多语教学老师
- 提示多样性与回应流畅性是关键,适合资源少语言
用语言模型生成有监督微调数据来训练小模型完成多语任务越来越普遍,但教师模型选择往往随意,常直接选用最大模型,而这些模型在非英语语言上可能能力不足。这会导致合成数据质量差,影响学生模型性能。本文系统分析有效多语教师的特征,结合内在数据质量与外在学生模型表现,提出‘多语得分’(Polyglot Score)评估指标。评估了10个语言模型在6种类型多样语言上的表现,生成超140万条SFT样本,训练240个学生模型。结果发现:模型规模不能有效预测教师效能;最有效的教师均小于最大模型,且其排名在不同学生模型家族间稳定。数据质量如提示多样性、长度和回应流畅性可解释93.3%的内在数据质量方差,并准确预测学生表现。最后提供实用建议:教师与学生模型应同家族,优先使用现有提示或英译生成,这对低资源语言尤其有益。希望推动多语合成数据与语言模型发展的数据驱动研究。
原文摘要 · Abstract (English)
Synthesizing supervised finetuning (SFT) data from language models (LMs) to teach smaller models multilingual tasks has become increasingly common. However, teacher model selection is often ad hoc, typically defaulting to the largest available option, even though such models may have significant capability gaps in non-English languages. This practice can result in poor-quality synthetic data and suboptimal student downstream performance. In this work, we systematically characterize what makes an effective multilingual teacher. We combine intrinsic measures of data quality with extrinsic student model performance in a metric we call Polyglot Score. We evaluate 10 LMs across 6 typologically diverse languages, generating over 1.4M SFT examples and training 240 student models. Our analyses reveal that model scale alone does not significantly predict teacher effectiveness: the most effective teachers we identify are consistently smaller than the largest models evaluated, and their ranking is stable across student base model families. Instead, data qualities such as prompt diversity, length, and response fluency capture 93.3% of the variance in intrinsic data quality and predict student performance. Finally, we provide practical recommendations, including matching the model families of teacher-student pairs and generating responses to existing prompts or translating them from English, which can yield improvements for less-resourced languages. We hope that our work advances data-centric research in multilingual synthetic data and LM development.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。