arXiv:2412.04261cs.CL2024-12被引 135

Aya Expanse 8B/32B多语言模型超越同类,32B版胜过70B的Llama 3.1。

Aya Expanse: Combining Research Breakthroughs for a New Multilingual Frontier

  • 融合数据精选、多语言偏好训练与模型合并技术提升性能。
  • 在23种语言的m-ArenaHard测试中,32B版赢率54.0%,超Llama 3.1 70B。
  • 开源权重并发布新多语言评测集,适合多语言研究者使用。

我们推出Aya Expanse模型家族,包含80亿和320亿参数的多语言大模型,旨在解决高性能多语言模型开发难题,使其能力媲美甚至超越单语模型。基于Cohere For AI与Cohere多年研究成果,包括数据套利、多语言偏好训练和模型合并等技术,Aya Expanse在多语言性能上达到新基准。在翻译成23种语言的Arena-Hard-Auto数据集上评估显示,Aya Expanse 8B和32B分别优于同参数规模的Gemma 2、Qwen 2.5和Llama 3.1,最高胜率达到76.6%。尤其值得注意的是,Aya Expanse 32B在与参数量两倍的Llama 3.1 70B对比中取得54.0%的胜率。本文还提供了Aya Expanse系列的扩展评估结果,并开放其模型权重及新发布的多语言评测集m-ArenaHard。

原文摘要 · Abstract (English)

We introduce the Aya Expanse model family, a new generation of 8B and 32B parameter multilingual language models, aiming to address the critical challenge of developing highly performant multilingual models that match or surpass the capabilities of monolingual models. By leveraging several years of research at Cohere For AI and Cohere, including advancements in data arbitrage, multilingual preference training, and model merging, Aya Expanse sets a new state-of-the-art in multilingual performance. Our evaluations on the Arena-Hard-Auto dataset, translated into 23 languages, demonstrate that Aya Expanse 8B and 32B outperform leading open-weight models in their respective parameter classes, including Gemma 2, Qwen 2.5, and Llama 3.1, achieving up to a 76.6% win-rate. Notably, Aya Expanse 32B outperforms Llama 3.1 70B, a model with twice as many parameters, achieving a 54.0% win-rate. In this short technical report, we present extended evaluation results for the Aya Expanse model family and release their open-weights, together with a new multilingual evaluation dataset m-ArenaHard.

多语言模型模型性能开源模型评测数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。