arXiv:2602.02287cs.CL2026-02被引 1

控制生成条件后,大模型评价在芬兰-乌戈尔语族中稳定性差,暴露了跨语言评估的可靠性问题。

Cross-Lingual Stability of LLM Judges Under Controlled Generation: Evidence from Finno-Ugric Languages

  • 固定生成参数,对比爱沙尼亚语、芬兰语、匈牙利语下的模型表现
  • 表面指标稳定,但语用判断(连贯性、指令遵循)排名严重颠倒
  • 揭示零样本评价在形态丰富语言中不可靠,需针对语言校准

大语言模型跨语言评估常混淆真实性能差异与测量不稳定性。我们通过控制生成条件,仅改变目标语言,研究评估可靠性。使用相同参数生成爱沙尼亚语、芬兰语和匈牙利语的合成客服对话,测试自动指标与大模型作为裁判的评分是否在三种形态丰富的相关语言间保持一致。以少量爱沙尼亚语母语者标注为基准,发现系统性排名不稳定:表层指标(词汇多样性、表面与语义相似性)具跨语言稳定性,但语用判断(连贯性、指令遵循)出现排名反转,相关性接近零。由于生成条件一致,此类不一致反映的是裁判评分在不同语言中的行为差异,而非模型真实性能。该受控设计提供诊断工具:若评估方法在相同生成下无法保持稳定,说明部署前已存在迁移失败风险。研究提示,在形态丰富语言中,零样本裁判迁移对话语级评估不可靠,应基于目标语言的人类基线进行语言特异性校准。数据、协议与评估框架已在 https://github.com/isaac-chung/cross-lingual-stability-judges 开放。

原文摘要 · Abstract (English)

Cross-lingual evaluation of large language models (LLMs) typically conflates two sources of variance: genuine model performance differences and measurement instability. We investigate evaluation reliability by holding generation conditions constant while varying target language. Using synthetic customer-support dialogues generated with identical parameters across Estonian, Finnish, and Hungarian, we test whether automatic metrics and LLM-as-a-judge scoring produce stable model rankings across these morphologically rich, related Finno-Ugric languages. With a small set of Estonian native speaker annotations as a reference point, we find systematic ranking instabilities: surface-level metrics (lexical diversity, surface and semantic similarity) maintain cross-language stability, but pragmatic judgments (coherence, instruction-following) exhibit rank inversions and near-zero correlations. Because generation is controlled, these inconsistencies reflect how judge scoring behaves differently across languages rather than true model differences. This controlled design provides a diagnostic probe: evaluation methods that fail to maintain stability under identical generation conditions signal transfer failure before deployment. Our findings suggest that zero-shot judge transfer is unreliable for discourse-level assessment in morphologically rich languages, motivating language-specific calibration against targeted human baselines. We release our controlled generation protocol, synthetic data, and evaluation framework to enable replication across language families at https://github.com/isaac-chung/cross-lingual-stability-judges.

跨语言评估模型评判形态丰富语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。