arXiv:2603.18557cs.CL2026-03被引 3

跨语言大模型评估新框架,无需标注就能迁移。

Cross-Lingual LLM-Judge Transfer via Evaluation Decomposition

  • 拆解评估维度构建通用标准,实现跨语言可迁移
  • 多语言任务上优于强基线,且无需目标语言标注
  • 适合需要低成本多语言评测的AI研究者

随着大语言模型在各类实际应用中日益普及,将自动化评估拓展至非英语语种已成为关键挑战。现有评估方法主要聚焦英语,而将其适配到其他语言受限于多数语言中高质量人工标注数据的稀缺与高昂成本。本文提出一种基于分解的评估框架,核心为通用评价维度集(UCS),该集合包含一组共享且语言无关的评估维度,生成可解释的中间表示,支持低监督下的跨语言迁移。在多种语言和模型架构上的忠实性任务实验表明,该方法持续优于强基线,且无需目标语言的人工标注。

原文摘要 · Abstract (English)

As large language models are increasingly deployed across diverse real-world applications, extending automated evaluation beyond English has become a critical challenge. Existing evaluation approaches are predominantly English-focused, and adapting them to other languages is hindered by the scarcity and cost of human-annotated judgments in most languages. We introduce a decomposition-based evaluation framework built around a Universal Criteria Set (UCS). UCS consists of a shared, language-agnostic set of evaluation dimensions, producing an interpretable intermediate representation that supports cross-lingual transfer with minimal supervision. Experiments on multiple faithfulness tasks across languages and model backbones demonstrate consistent improvements over strong baselines without requiring target-language annotations.

大模型评估跨语言无监督

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。