首次评估多语言大模型跨语言遗忘效果,发现多数方法无效。
Evaluating Cross-Lingual Unlearning in Multilingual Language Models
- 用七种语言变体的翻译版TOFU基准测试遗忘算法
- 子空间投影法在跨语言遗忘中表现最佳,性能下降小
- 多语言模型存在共享语义结构,适合用子空间方法遗忘
我们首次全面评估了多语言大模型中的跨语言遗忘。基于七种语言/文字变体的翻译版TOFU基准,测试了主流遗忘算法,发现大多数方法无法在保持模型性能的同时,有效移除非训练语言中的事实信息。然而,子空间投影法始终表现更优,能在最小性能损失下实现强跨语言遗忘。对学习到的任务子空间的分析显示存在共享的跨语言结构:移除该共享子空间会影响所有语言,而移除特定语言子空间则仅影响对应语言。结果表明,多语言遗忘依赖于权重空间中的几何结构,为未来遗忘系统设计提供了子空间方法的新方向。
原文摘要 · Abstract (English)
We present the first comprehensive evaluation of cross-lingual unlearning in multilingual LLMs. Using translated TOFU benchmarks in seven language/script variants, we test major unlearning algorithms and show that most fail to remove facts outside the training language, even when utility remains high. However, subspace-projection consistently outperforms the other methods, achieving strong cross-lingual forgetting with minimal degradation. Analysis of learned task subspaces reveals a shared interlingua structure: removing this shared subspace harms all languages, while removing language-specific components selectively affects one. These results demonstrate that multilingual forgetting depends on geometry in weight space, motivating subspace-based approaches for future unlearning systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。