arXiv:2510.07877cs.CL2025-10ACL被引 1

评测多语言大模型翻译质量与偏见,发现低资源语言表现差且易放大偏见

Ready to Translate, Not to Represent? Bias and Performance Gaps in Multilingual LLMs Across Language Families and Domains

  • 构建统一框架Translation Tangles,跨语言对与领域评估24组双向翻译
  • 在1439对翻译-参考对中发现低资源语言偏差显著,性能差距达30%以上
  • 提出混合检测方法,结合规则、语义相似度与大模型验证,可识别隐性偏见

大型语言模型(LLMs)推动机器翻译发展,实现数百种语言和文本领域的上下文感知、流畅翻译。然而,其在不同语言家族和专业领域间表现不均,且可能在训练数据中编码并放大各类偏见,尤其对低资源语言带来公平性风险。为此,本文提出Translation Tangles框架与数据集,用于统一评估开源LLMs的翻译质量与公平性。我们对24个双向语言对在多个领域进行基准测试,采用多种指标评估。进一步设计了一种融合规则启发式、语义相似度过滤与大模型验证的混合偏见检测流程,并基于1,439对翻译-参考对的人工评估构建高质量标注数据集。代码与数据已公开于GitHub。

原文摘要 · Abstract (English)

The rise of Large Language Models (LLMs) has redefined Machine Translation (MT), enabling context-aware and fluent translations across hundreds of languages and textual domains. Despite their remarkable capabilities, LLMs often exhibit uneven performance across language families and specialized domains. Moreover, recent evidence reveals that these models can encode and amplify different biases present in their training data, posing serious concerns for fairness, especially in low-resource languages. To address these gaps, we introduce Translation Tangles, a unified framework and dataset for evaluating the translation quality and fairness of open-source LLMs. Our approach benchmarks 24 bidirectional language pairs across multiple domains using different metrics. We further propose a hybrid bias detection pipeline that integrates rule-based heuristics, semantic similarity filtering, and LLM-based validation. We also introduce a high-quality, bias-annotated dataset based on human evaluations of 1,439 translation-reference pairs. The code and dataset are accessible on GitHub: https://github.com/faiyazabdullah/TranslationTangles

大模型机器翻译偏见检测公平性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。