arXiv:2603.17952cs.CL2026-03

检测解码器模型的性别偏见,发现指令微调能降低男性默认倾向。

Gender Disambiguation in Machine Translation: Diagnostic Evaluation in Decoder-Only Architectures

  • 提出新指标'先验偏见',衡量模型对性别的默认假设。
  • 解码器模型在性别标注任务上表现不如编码器-解码器架构。
  • 指令微调可提升上下文理解并减少男性化倾向。

尽管大语言模型在众多自然语言任务中表现优异,但仍存在系统性偏见。其中,由于语言间性别标记方式差异显著,机器翻译中的性别偏见尤为突出。翻译常需将源语言中的隐含性别信号转化为显式标记形式。现有基准测试虽能反映整体差距,但难以捕捉现代机器翻译中性别偏见的复杂性。本文扩展了偏见评估框架:(i) 提出新指标‘先验偏见’,量化模型对性别的默认倾向;(ii) 将该框架应用于解码器仅架构的机器翻译模型。结果表明,尽管解码器模型规模大且性能领先,但在性别相关指标上并未普遍优于编码器-解码器架构;然而,后训练(如指令微调)不仅能增强上下文感知能力,还能有效降低男性的先验偏见。

原文摘要 · Abstract (English)

While Large Language Models achieve state-of-the-art results across a wide range of NLP tasks, they remain prone to systematic biases. Among these, gender bias is particularly salient in MT, due to systematic differences across languages in whether and how gender is marked. As a result, translation often requires disambiguating implicit source signals into explicit gender-marked forms. In this context, standard benchmarks may capture broad disparities but fail to reflect the full complexity of gender bias in modern MT. In this paper, we extend recent frameworks on bias evaluation by: (i) introducing a novel measure coined "Prior Bias", capturing a model's default gender assumptions, and (ii) applying the framework to decoder-only MT models. Our results show that, despite their scale and state-of-the-art status, decoder-only models do not generally outperform encoder-decoder architectures on gender-specific metrics; however, post-training (e.g., instruction tuning) not only improves contextual awareness but also reduces the masculine Prior Bias.

性别偏见机器翻译模型评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。