arXiv:2503.01996cs.CL2025-03被引 22

首个跨26语言的长文本模型评测基准,发现英语并非最强。

One ruler to measure them all: Benchmarking multilingual long-context language models

  • 构建26语言长文本评测集,含7类合成任务
  • 128K上下文下低资源语言性能差距扩大
  • 跨语言指令时性能波动达20%,适合多语研究者

我们提出ONERULER,一个涵盖26种语言的多语言长上下文语言模型评测基准。该基准在英文RULER的基础上扩展,包含7个合成任务,测试检索与聚合能力,并引入无答案场景的'针在草堆'变体。通过先撰写英文指令再由母语者翻译至25种语言的两步流程构建。实验显示,当上下文长度从8K增至128K时,低资源语言与高资源语言间性能差距扩大;出人意料的是,英语在26种语言中仅排名第6,波兰语表现最佳。此外,许多模型(尤其是OpenAI的o3-mini-high)即使在高资源语言中也错误预测答案缺失。在跨语言场景下(指令与上下文语言不同),性能波动最高可达20%。ONERULER的发布旨在推动多语言与跨语言长上下文训练方法的研究。

原文摘要 · Abstract (English)

We present ONERULER, a multilingual benchmark designed to evaluate long-context language models across 26 languages. ONERULER adapts the English-only RULER benchmark (Hsieh et al., 2024) by including seven synthetic tasks that test both retrieval and aggregation, including new variations of the "needle-in-a-haystack" task that allow for the possibility of a nonexistent needle. We create ONERULER through a two-step process, first writing English instructions for each task and then collaborating with native speakers to translate them into 25 additional languages. Experiments with both open-weight and closed LLMs reveal a widening performance gap between low- and high-resource languages as context length increases from 8K to 128K tokens. Surprisingly, English is not the top-performing language on long-context tasks (ranked 6th out of 26), with Polish emerging as the top language. Our experiments also show that many LLMs (particularly OpenAI's o3-mini-high) incorrectly predict the absence of an answer, even in high-resource languages. Finally, in cross-lingual scenarios where instructions and context appear in different languages, performance can fluctuate by up to 20% depending on the instruction language. We hope the release of ONERULER will facilitate future research into improving multilingual and cross-lingual long-context training pipelines.

多语言长上下文评测基准跨语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。