arXiv:2606.08092cs.CL2026-06被引 1

让不同语言的评价分歧成为改进多语言模型判断力的工具

When Languages Disagree: Self-Evolving Multilingual LLM Judges

论文配图:When Languages Disagree: Self-Evolving Multilingual LLM Judges
图 1 · 摘自论文原文
  • 通过收集多语言独立评价并反馈不一致结果,实现自我迭代优化
  • 在多个基准上表现优于投票和单语言基准,提升准确率与跨语言一致性
  • 适合需要高鲁棒性多语言评估的场景,如跨国AI系统评测

多语言大模型作为评判者被广泛用于跨语言输出评估,但存在跨语言不一致问题(Fu and Liu, 2025)。现有方法通常将这种不一致视为噪声并用投票或聚合缓解。本文发现,跨语言不一致反而可提供互补评估信号。我们的最优分析表明,跨语言采样判断的性能上限高于单一语言判断,说明不同语言可能包含互补判断。受此启发,我们提出SEMJ:一种自演化多语言评判框架,利用跨语言不一致进行迭代优化。SEMJ构建每个输入的多语言变体,收集独立判断与推理过程,并将不一致输出反馈用于自我反思与重评。实验在多个基准上显示,SEMJ在准确率与跨语言一致性上持续优于投票与反思基线。进一步分析表明,不一致触发了有效重评,提升了判断质量。

原文摘要 · Abstract (English)

Multilingual LLM-as-a-judge is widely used to evaluate model outputs across languages, but suffers from cross-lingual inconsistency (Fu and Liu, 2025). Existing methods typically treat this inconsistency as noise and mitigate it through voting or aggregation. In this work, we instead show that multilingual inconsistency can provide complementary evaluation signals. Our oracle analysis finds that sampling judgments across languages yields a higher performance upper bound than single-language judging, indicating that different languages potentially include complementary judgments. Motivated by this finding, we propose SEMJ, a self-evolving multilingual judge that leverages cross-lingual inconsistency for iterative refinement. SEMJ constructs multilingual variants of each input, collects independent judgments and rationales, and feeds inconsistent outputs back for self-reflection and re-evaluation. Experiments on multiple benchmarks show that SEMJ consistently outperforms voting and reflection baselines in both accuracy and cross-lingual consistency. Further analysis shows that inconsistency triggers useful re-evaluation, which improves judgment quality.

多语言评估自演化模型评判一致性优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。