arXiv:2509.04032cs.CLcs.LG2025-09中稿 · to Multilingual Re…被引 1

用新指标测20种语言模型表现,发现大模型跨语言一致性更强。

What if I ask in \textit{alia lingua}? Measuring Functional Similarity Across Languages

  • 用κ_p指标量化20语言47主题的模型输出相似性
  • 模型越大越一致,跨语言表现随能力提升而增强
  • 模型自身跨语言一致性高于不同模型同语种对比

本文使用新提出的模型相似性度量κ_p,分析全球多语言测评集GlobalMMLU中20种语言、47个主题下的模型输出。结果表明,随着模型规模和能力增大,其跨语言响应的一致性显著提升。有趣的是,模型自身的跨语言一致性高于不同模型在同一种语言下的对齐程度。这不仅证明了κ_p在评估多语言可靠性方面的实用性,也为构建更一致的多语言系统提供了指导方向。

原文摘要 · Abstract (English)

How similar are model outputs across languages? In this work, we study this question using a recently proposed model similarity metric $κ_p$ applied to 20 languages and 47 subjects in GlobalMMLU. Our analysis reveals that a model's responses become increasingly consistent across languages as its size and capability grow. Interestingly, models exhibit greater cross-lingual consistency within themselves than agreement with other models prompted in the same language. These results highlight not only the value of $κ_p$ as a practical tool for evaluating multilingual reliability, but also its potential to guide the development of more consistent multilingual systems.

多语言模型评估一致性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。