FairQE通过多智能体机制减少翻译质量评估中的性别偏见。
FairQE: Multi-Agent Framework for Mitigating Gender Bias in Translation Quality Estimation

- 用多智能体检测性别线索并生成性别反转译文变体。
- 在多种场景下显著提升性别公平性,且不降低评估准确性。
- 适合关注公平性与可靠性的机器翻译评估研究者使用。
质量估计(QE)旨在无参考译文的情况下评估机器翻译质量,但近期研究表明现有QE模型存在系统性性别偏见:在性别模糊语境中倾向于偏好男性表达,甚至在性别明确时仍给性别错配译文打高分。为此,我们提出FairQE,一种基于多智能体的公平性感知QE框架,可缓解性别模糊和性别明确场景下的性别偏见。FairQE检测性别线索,生成性别翻转的译文变体,并通过动态偏差感知聚合机制,结合传统QE得分与大模型驱动的偏见缓解推理。该设计在保持原有QE模型优势的同时,以即插即用方式校准其性别相关偏见。跨多个性别偏见评估设置的大量实验表明,FairQE在公平性上持续优于强基线。此外,在遵循WMT 2023评测任务标准的MQM元评估中,FairQE达到具有竞争力或更优的通用QE性能。结果表明,可在不牺牲评估准确性的前提下有效缓解QE中的性别偏见,实现更公平、可靠的翻译评估。
原文摘要 · Abstract (English)
Quality Estimation (QE) aims to assess machine translation quality without reference translations, but recent studies have shown that existing QE models exhibit systematic gender bias. In particular, they tend to favor masculine realizations in gender-ambiguous contexts and may assign higher scores to gender-misaligned translations even when gender is explicitly specified. To address these issues, we propose FairQE, a multi-agent-based, fairness-aware QE framework that mitigates gender bias in both gender-ambiguous and gender-explicit scenarios. FairQE detects gender cues, generates gender-flipped translation variants, and combines conventional QE scores with LLM-based bias-mitigating reasoning through a dynamic bias-aware aggregation mechanism. This design preserves the strengths of existing QE models while calibrating their gender-related biases in a plug-and-play manner. Extensive experiments across multiple gender bias evaluation settings demonstrate that FairQE consistently improves gender fairness over strong QE baselines. Moreover, under MQM-based meta-evaluation following the WMT 2023 Metrics Shared Task, FairQE achieves competitive or improved general QE performance. These results show that gender bias in QE can be effectively mitigated without sacrificing evaluation accuracy, enabling fairer and more reliable translation evaluation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。