arXiv:2512.13298cs.CLcs.AI2025-12

10亿参数的多语言模型,支持13种欧洲语言,性能超越同类小模型。

MiniLingua: A Small Open-Source LLM for European Languages

  • 从零训练10亿参数模型,专注13种欧洲语言
  • 在摘要、分类等任务上优于更大规模的EuroLLM
  • 开源权重与代码,适合本地部署与多语言研究

大型语言模型虽强大,但存在计算成本高、隐私风险及英语中心化等问题。近期研究表明,约10亿参数的小型高效模型可实现良好性能并支持设备端运行。本文提出MiniLingua,一个从零训练的10亿参数多语言大模型,专为13种欧洲语言设计,兼顾语言覆盖与指令遵循能力。评估显示,其指令微调版本在摘要生成、分类及开闭卷问答任务上优于采用相似训练策略但训练预算更大的EuroLLM,在开放生成任务上也保持与更先进模型相当的竞争力。本文公开模型权重、分词器及数据处理与训练源码。

原文摘要 · Abstract (English)

Large language models are powerful but often limited by high computational cost, privacy concerns, and English-centric training. Recent progress demonstrates that small, efficient models with around one billion parameters can deliver strong results and enable on-device use. This paper introduces MiniLingua, a multilingual open-source LLM of one billion parameters trained from scratch for 13 European languages, designed to balance coverage and instruction-following capabilities. Based on evaluation results, the instruction-tuned version of MiniLingua outperforms EuroLLM, a model with a similar training approach but a larger training budget, on summarization, classification and both open- and closed-book question answering. Moreover, it remains competitive with more advanced state-of-the-art models on open-ended generation tasks. We release model weights, tokenizer and source code used for data processing and model training.

小模型多语言开源欧洲语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。