在跨语言隐空间中去偏,提升多语言大模型的公平性与迁移能力。
Debiasing Multilingual LLMs in Cross-lingual Latent Space
- 在对齐的跨语言隐空间中进行去偏,而非直接作用于模型表征。
- 在英、法、德、荷四语言上,去偏效果和跨语言迁移性均显著提升。
- 适合关注多语言模型公平性与跨语言泛化能力的研究者。
去偏技术如SentDebias旨在降低大语言模型中的偏见。以往研究通过直接在大模型表示上应用这些方法评估其跨语言可迁移性,发现效果有限。本文提出在联合隐空间中进行去偏,而非直接作用于大模型表示。我们使用平行TED演讲脚本训练自编码器,构建了一个对齐良好的跨语言隐空间。在Aya-expanse模型及两种去偏技术下,针对英语、法语、德语、荷兰语四个语言的实验表明:(a) 自编码器能有效构建对齐的跨语言隐空间;(b) 在学习到的跨语言隐空间中应用去偏技术,显著提升了整体去偏性能与跨语言可迁移性。
原文摘要 · Abstract (English)
Debiasing techniques such as SentDebias aim to reduce bias in large language models (LLMs). Previous studies have evaluated their cross-lingual transferability by directly applying these methods to LLM representations, revealing their limited effectiveness across languages. In this work, we therefore propose to perform debiasing in a joint latent space rather than directly on LLM representations. We construct a well-aligned cross-lingual latent space using an autoencoder trained on parallel TED talk scripts. Our experiments with Aya-expanse and two debiasing techniques across four languages (English, French, German, Dutch) demonstrate that a) autoencoders effectively construct a well-aligned cross-lingual latent space, and b) applying debiasing techniques in the learned cross-lingual latent space significantly improves both the overall debiasing performance and cross-lingual transferability.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。