为罗马尼亚语打造的小型高效基础模型,性能媲美大型开源模型
LLMic: Romanian Foundation Language Model
- 针对低资源语言罗马尼亚语构建完整预训练流程
- 微调后在英罗翻译任务上超越现有方案
- 适合关注小语种AI落地的研究者与开发者
近年来大语言模型在各类任务中表现出色,商业模型引领发展。尽管开源模型规模较小,但通过专业化和微调仍具竞争力。然而,开源模型在低资源语言上表现不佳,因训练数据覆盖不足。本文提出 LLMic,一个专为罗马尼亚语设计的双语基础语言模型。我们完整记录了低资源语言基础模型的预训练过程,包括语料构建、架构选择与超参数优化。评估表明,LLMic 可有效适配目标语言任务,性能可比肩更大规模的开源模型。在初始预训练后对翻译任务进行微调,其在英罗翻译任务中表现优于现有方案。该成果为罗马尼亚语社区提供了高效的大规模语言处理能力,仅需更小的 LLMic 模型即可实现。
原文摘要 · Abstract (English)
Recent advances in Large Language Models (LLMs) have demonstrated remarkable capabilities across various tasks with commercial models leading the way. While open models usually operate at a smaller scale, they maintain competitiveness through specialization and fine-tuning. However, a significant challenge persists: open models often underperform in low-resource languages due to limited representation in the training corpus. In this paper, we present LLMic, a bilingual foundation language model designed specifically for the Romanian Language. We document the complete process of pretraining a foundation model for a low-resource language, including corpus construction, architecture selection, and hyper-parameter optimization. Our evaluation demonstrates that LLMic can be specialized for tasks in the target language, achieving results comparable to other much larger open models. We show that fine-tuning LLMic for language translation after the initial pretraining phase outperforms existing solutions in English-to-Romanian translation tasks. This opens the path for efficient large-scale processing for the Romanian language community, using the much smaller LLMic model
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。