arXiv:2505.21999cs.CL2025-05被引 2

用翻译后评估法检测大模型跨语言一致性,发现多语言表现严重不稳。

Found in Translation: Measuring Multilingual LLM Consistency as Simple as Translate then Evaluate

  • 通过翻译再评估策略,量化模型在不同语言间的响应一致性。
  • 在30种语言中发现显著不一致,部分语系和文字系统性能大幅下降。
  • 提供可复现的评估框架,适合未来多语言模型基准测试。

大型语言模型在英文问答中表现优异,但其在其他语言中的响应是否一致仍存疑。现有评测依赖昂贵标注数据,且对开放式生成任务(存在多个正确答案)难以衡量。本文提出一种基于‘翻译后评估’的跨语言一致性评测框架,从信息与共情两个维度进行验证。实验覆盖30种语言,结果显示主流大模型在跨语言表现上存在显著不一致,特定语系与书写系统中性能严重下滑,暴露出其多语言能力的关键短板。研究呼吁开展多维度、一致性的跨语言评测。我们邀请从业者使用该框架进行未来多语言模型基准测试。

原文摘要 · Abstract (English)

Large language models (LLMs) provide detailed and impressive responses to queries in English. However, are they really consistent at responding to the same query in other languages? The popular way of evaluating for multilingual performance of LLMs requires expensive-to-collect annotated datasets. Further, evaluating for tasks like open-ended generation, where multiple correct answers may exist, is nontrivial. Instead, we propose to evaluate the predictability of model response across different languages. In this work, we propose a framework to evaluate LLM's cross-lingual consistency based on a simple Translate then Evaluate strategy. We instantiate this evaluation framework along two dimensions of consistency: information and empathy. Our results reveal pronounced inconsistencies in popular LLM responses across thirty languages, with severe performance deficits in certain language families and scripts, underscoring critical weaknesses in their multilingual capabilities. These findings necessitate cross-lingual evaluations that are consistent along multiple dimensions. We invite practitioners to use our framework for future multilingual LLM benchmarking.

多语言一致性评测框架

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。