AMALIA是首个专注葡萄牙语的开源大模型,提升本地语言表现。
AMALIA Technical Report: A Fully Open Source Large Language Model for European Portuguese
- 用高质量葡萄牙语数据进行中后期训练,强化本地语言能力。
- 在本土化评测中显著优于基线模型,尤其在生成与语义理解上。
- 适合关注欧洲葡语技术落地的研究者与开发者使用。
尽管开源大语言模型进展迅速,欧洲葡萄牙语(pt-PT)在训练数据和原生评估中仍严重不足,机器翻译的基准可能遗漏其语言与文化特征。我们提出AMALIA,一个完全开源的大语言模型,通过在中后期训练阶段使用更多高质量pt-PT数据,优先支持pt-PT。为更真实评估pt-PT能力,我们发布一套包含翻译标准任务及四个新数据集的评测套件,涵盖pt-PT生成、语言能力与pt-PT/pt-BR偏见检测。实验表明,AMALIA在翻译基准上匹配强基线,在pt-PT专属评测中表现显著提升,支持针对欧洲葡萄牙语进行定向训练与原生评测。
原文摘要 · Abstract (English)
Despite rapid progress in open large language models (LLMs), European Portuguese (pt-PT) remains underrepresented in both training data and native evaluation, with machine-translated benchmarks likely missing the variant's linguistic and cultural nuances. We introduce AMALIA, a fully open LLM that prioritizes pt-PT by using more high-quality pt-PT data during both the mid- and post-training stages. To evaluate pt-PT more faithfully, we release a suite of pt-PT benchmarks that includes translated standard tasks and four new datasets targeting pt-PT generation, linguistic competence, and pt-PT/pt-BR bias. Experiments show that AMALIA matches strong baselines on translated benchmarks while substantially improving performance on pt-PT-specific evaluations, supporting the case for targeted training and native benchmarking for European Portuguese.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。