跨语言去偏方法,用迭代奇异值分解消除多语种共现偏见
Iterative Multilingual Spectral Attribute Erasure
- 通过迭代SVD截断识别多语言联合偏见子空间
- 在8种语言、5个性别维度上验证,零样本场景也有效
- 适配主流大模型,兼顾去偏效果与模型性能
多语言表示将语义相近的词映射到统一语义空间,为跨语言去偏迁移提供了可能。然而,现有去偏方法仅针对单一语言,无法利用这一优势。我们提出迭代多语言谱属性擦除(IMSAE),通过迭代基于奇异值分解(SVD)的截断,识别并缓解多语言间的联合偏见子空间。在8种语言和5个社会人口学维度上评估,IMSAE在标准设置和零样本设置(目标语言数据缺失,但可借助语言相似语言去偏)中均表现优异。对BERT、LLaMA、Mistral等多种语言模型的全面实验表明,IMSAE优于传统单语言与跨语言方法,同时保持模型实用性。
原文摘要 · Abstract (English)
Multilingual representations embed words with similar meanings to share a common semantic space across languages, creating opportunities to transfer debiasing effects between languages. However, existing methods for debiasing are unable to exploit this opportunity because they operate on individual languages. We present Iterative Multilingual Spectral Attribute Erasure (IMSAE), which identifies and mitigates joint bias subspaces across multiple languages through iterative SVD-based truncation. Evaluating IMSAE across eight languages and five demographic dimensions, we demonstrate its effectiveness in both standard and zero-shot settings, where target language data is unavailable, but linguistically similar languages can be used for debiasing. Our comprehensive experiments across diverse language models (BERT, LLaMA, Mistral) show that IMSAE outperforms traditional monolingual and cross-lingual approaches while maintaining model utility.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。