arXiv:2603.11510cs.CL2026-03被引 12

33.5亿参数模型实现多语言顶尖表现,兼顾效率与平衡性。

Tiny Aya: Bridging Scale and Multilingual Depth

  • 基于70种语言训练,通过区域感知微调提升多语言能力
  • 仅3.35亿参数即达翻译质量与生成效果的顶尖水平
  • 适合需要高效部署的多语言应用开发者

Tiny Aya重新定义了小型多语言模型的潜力。在70种语言上训练,并通过区域感知后训练优化,仅用3.35亿参数即实现了顶尖的翻译质量、强大的多语言理解能力以及高质量的目标语言生成。发布的模型包括一个预训练基础模型、一个全球均衡指令微调版本,以及三个针对非洲、南亚、欧洲、亚太和西亚地区的区域专精模型。本文详述了Tiny Aya的训练策略、数据构成及全面评估框架,提出了一条以效率、语言间性能均衡和实际部署为导向的多语言AI新路径。

原文摘要 · Abstract (English)

Tiny Aya redefines what a small multilingual language model can achieve. Trained on 70 languages and refined through region-aware posttraining, it delivers state-of-the-art in translation quality, strong multilingual understanding, and high-quality target-language generation, all with just 3.35B parameters. The release includes a pretrained foundation model, a globally balanced instruction-tuned variant, and three region-specialized models targeting languages from Africa, South Asia, Europe, Asia-Pacific, and West Asia. This report details the training strategy, data composition, and comprehensive evaluation framework behind Tiny Aya, and presents an alternative scaling path for multilingual AI: one centered on efficiency, balanced performance across languages, and practical deployment.

多语言模型小模型高效生成跨语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。