语言模型习得成语表达慢,但翻译指令微调后很快遗忘。
Preferences for Idiomatic Language are Acquired Slowly -- and Forgotten Quickly: A Case Study on Swedish
- 通过最小差异对测试模型在瑞典语中对成语与语法正确性的偏好
- 大模型(80亿参数)在成语能力上持续提升,训练越久越好
- 用英译数据进行指令微调会迅速削弱模型的成语偏好
本研究探究语言模型在预训练及从英语迁移到瑞典语过程中,对成语性表达与语法可接受性表达的偏好形成机制。我们从零开始训练瑞典语模型,并微调英语预训练模型,在不同检查点使用最小差异对进行探测。针对语法可接受性,将现有基准改编为最小差异对格式;为评估成语性,提出两个新数据集:一个对比固定成语与其合理变体,另一个对比成语化瑞典语与翻译腔表达。结果表明,成语能力的获得比语法和词汇正确性更缓慢,且随着训练时间延长,多数任务性能趋于饱和,但成语相关表现仍在持续提升,尤其在最大模型(80亿参数)中明显。然而,使用机器翻译的英文指令数据进行指令微调,会导致模型快速丧失对成语表达的偏好。
原文摘要 · Abstract (English)
In this study, we investigate how language models develop preferences for \textit{idiomatic} as compared to \textit{linguistically acceptable} Swedish, both during pretraining and when adapting a model from English to Swedish. To do so, we train models on Swedish from scratch and by fine-tuning English-pretrained models, probing their preferences at various checkpoints using minimal pairs that differ in linguistic acceptability or idiomaticity. For linguistic acceptability, we adapt existing benchmarks into a minimal-pair format. To assess idiomaticity, we introduce two novel datasets: one contrasting conventionalized idioms with plausible variants, and another contrasting idiomatic Swedish with Translationese. Our findings suggest that idiomatic competence emerges more slowly than other linguistic abilities, including grammatical and lexical correctness. While longer training yields diminishing returns for most tasks, idiom-related performance continues to improve, particularly in the largest model tested (8B). However, instruction tuning on data machine-translated from English -- the common approach for languages with little or no native instruction data -- causes models to rapidly lose their preference for idiomatic language.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。