arXiv:2603.10351cs.CLcs.AI2026-03被引 1

提出新方法消除多语言大模型评分时对机器翻译文本的偏好

Mitigating Translationese Bias in Multilingual LLM-as-a-Judge via Disentangled Information Bottleneck

  • 通过分离关键判断特征与偏差因素来降低评分偏见
  • 在多语言评测中显著减少对机翻文本的系统性偏好
  • 适合需要公平评估多语言生成质量的研究者使用

大型语言模型(LLMs)已成为多语言评估的标准工具,但其存在严重的系统性翻译腔偏差:在低资源语言中,模型会系统性地偏好机器翻译文本而非人工撰写参考文本。本文将该偏差归因于(i)与英语潜在流形对齐的隐含关联,以及(ii)跨语言可预测性。为此,提出DIBJudge框架,通过变分信息压缩学习最小充分的判别性表征,同时将虚假因素显式分离至专用偏差分支。此外,引入交叉协方差惩罚项,主动抑制鲁棒表征与偏差表征间的统计依赖,促进有效解耦。在多语言奖励建模基准和专门设计的翻译腔偏差评测套件上,DIBJudge持续优于强基线,显著缓解了翻译腔偏差。

原文摘要 · Abstract (English)

Large language models (LLMs) have become a standard for multilingual evaluation, yet they exhibit a severe systematic translationese bias. In this paper, translationese bias is characterized as LLMs systematically favoring machine-translated text over human-authored references, particularly in low-resource languages. We attribute this bias to spurious correlations with (i) latent manifold alignment with English and (ii) cross-lingual predictability. To mitigate this bias, we propose DIBJudge, a robust fine-tuning framework that learns a minimally sufficient, judgment-critical representation via variational information compression, while explicitly isolating spurious factors into the dedicated bias branch. Furthermore, we incorporate a cross-covariance penalty that explicitly suppresses statistical dependence between robust and bias representations, thereby encouraging effective disentanglement. Extensive evaluations on multilingual reward modeling benchmarks and a dedicated translationese bias evaluation suite demonstrate that the proposed DIBJudge consistently outperforms strong baselines and substantially mitigates translationese bias.

大模型评估多语言偏差纠正信息瓶颈

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。