研究多语言大模型中删去知识的迁移规律,发现语法相似语言间效果更好。
Multilingual Amnesia: On the Transferability of Unlearning in Multilingual LLMs
- 在十种语言上测试删除知识的效果,涵盖高/低资源语言
- 高资源语言删去知识更稳定,语法相近语言间迁移更强
- 适合关注多语言模型安全与偏见控制的研究者
随着多语言大模型广泛应用,确保其在多元语言环境中的安全与公平面临独特挑战。现有机器遗忘研究多集中于单语场景(通常为英语),而多语言环境因跨语言知识迁移及预训练与微调数据中的偏见引入了额外复杂性。本文以Aya-Expanse 8B模型为基础,在数据遗忘与概念遗忘两种设置下开展研究,通过翻译将事实知识与刻板印象基准扩展至十种语言:英语、法语、阿拉伯语、日语、俄语、波斯语、韩语、印地语、希伯来语和印尼语,覆盖五大语系且资源水平差异显著。实验表明,高资源语言中遗忘更稳定,且在语言类型相近的语言间存在不对称迁移效应。对语言距离的分析显示,句法相似性是预测跨语言遗忘行为的最强因素。
原文摘要 · Abstract (English)
As multilingual large language models become more widely used, ensuring their safety and fairness across diverse linguistic contexts presents unique challenges. While existing research on machine unlearning has primarily focused on monolingual settings, typically English, multilingual environments introduce additional complexities due to cross-lingual knowledge transfer and biases embedded in both pretraining and fine-tuning data. In this work, we study multilingual unlearning using the Aya-Expanse 8B model under two settings: (1) data unlearning and (2) concept unlearning. We extend benchmarks for factual knowledge and stereotypes to ten languages through translation: English, French, Arabic, Japanese, Russian, Farsi, Korean, Hindi, Hebrew, and Indonesian. These languages span five language families and a wide range of resource levels. Our experiments show that unlearning in high-resource languages is generally more stable, with asymmetric transfer effects observed between typologically related languages. Furthermore, our analysis of linguistic distances indicates that syntactic similarity is the strongest predictor of cross-lingual unlearning behavior.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。