无需标注数据,用对话任务评估多语言生成能力。
MUG-Eval: A Proxy Evaluation Framework for Multilingual Generation Capabilities in Any Language
- 将现有基准转化为对话任务,以任务成功率衡量生成效果。
- 在30种语言上测试8个模型,相关性超0.75,结果可靠。
- 适合低资源语言评估,不依赖语言工具或大模型判别。
评估大语言模型(LLMs)的文本生成能力极具挑战,尤其对低资源语言而言,直接评估方法稀缺。我们提出MUG-Eval,一种新框架,通过将现有基准转换为对话任务来评估多语言生成能力,并测量模型在这些任务中的准确率。设计对话任务时特别强调目标语言的有效沟通。我们仅使用任务成功率为对话生成成功的代理指标。该方法具有两大优势:不依赖语言特定NLP工具或标注数据(对多数语言有限),也不依赖大模型作为评判者(其评估质量在少数高资源语言外显著下降)。我们在30种涵盖高、中、低资源水平的语言中评估了8个大模型,发现MUG-Eval与已有基准高度相关($r$ > 0.75),并实现跨语言、跨模型的标准化比较。该框架为多语言生成评估提供了一种稳健且高效的解决方案,可扩展至数千种语言。
原文摘要 · Abstract (English)
Evaluating text generation capabilities of large language models (LLMs) is challenging, particularly for low-resource languages where methods for direct assessment are scarce. We propose MUG-Eval, a novel framework that evaluates LLMs' multilingual generation capabilities by transforming existing benchmarks into conversational tasks and measuring the LLMs' accuracies on those tasks. We specifically designed these conversational tasks to require effective communication in the target language. Then, we simply use task success rate as a proxy for successful conversation generation. Our approach offers two key advantages: it is independent of language-specific NLP tools or annotated datasets, which are limited for most languages, and it does not rely on LLMs-as-judges, whose evaluation quality degrades outside a few high-resource languages. We evaluate 8 LLMs across 30 languages spanning high, mid, and low-resource categories, and we find that MUG-Eval correlates strongly with established benchmarks ($r$ > 0.75) while enabling standardized comparisons across languages and models. Our framework provides a robust and resource-efficient solution for evaluating multilingual generation that can be extended to thousands of languages.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。