arXiv:2604.02772cs.CL2026-04

多语言大模型全链路去偏方法,有效降低性别、种族、宗教偏见。

Multiple-Debias: A Full-process Debiasing Method for Multilingual Pre-trained Language Models

  • 融合跨语言反事实数据增强与自去偏机制,覆盖预处理到后处理全流程。
  • 在四种语言中显著降低三类敏感属性的偏见,优于单语去偏方法。
  • 适用于需要公平性的多语言NLP应用,如跨文化对话系统。

多语言预训练语言模型(MPLMs)已成为自然语言处理的重要工具,但常表现出与性别、种族、宗教等敏感属性相关的偏见。本文提出一种全链路多语言去偏方法 Multiple-Debias,通过在预处理和后处理阶段结合多语言反事实数据增强与多语言自去偏,并辅以参数高效微调,显著降低了四种语言中三类敏感属性的偏见。我们还将 CrowS-Pairs 扩展至德语、西班牙语、中文和日语,验证了该方法对性别、种族和宗教偏见的有效性。实验表明:(i) 多语言去偏方法在缓解偏见方面优于单语方法;(ii) 融合不同语言的去偏信息能显著提升 MPLMs 的公平性。

原文摘要 · Abstract (English)

Multilingual Pre-trained Language Models (MPLMs) have become essential tools for natural language processing. However, they often exhibit biases related to sensitive attributes such as gender, race, and religion. In this paper, we introduce a comprehensive multilingual debiasing method named Multiple-Debias to address these issues across multiple languages. By incorporating multilingual counterfactual data augmentation and multilingual Self-Debias across both pre-processing and post-processing stages, alongside parameter-efficient fine-tuning, we significantly reduced biases in MPLMs across three sensitive attributes in four languages. We also extended CrowS-Pairs to German, Spanish, Chinese, and Japanese, validating our full-process multilingual debiasing method for gender, racial, and religious bias. Our experiments show that (i) multilingual debiasing methods surpass monolingual approaches in effectively mitigating biases, and (ii) integrating debiasing information from different languages notably improves the fairness of MPLMs.

多语言去偏公平性预训练模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。