开源葡萄牙语大模型Tucano 2,性能领先且完全可复现。
Tucano 2 Cool: Better Open Source LLMs for Portuguese
- 基于新数据集GigaVerbo-v2及合成数据,构建多场景训练体系。
- 0.5-3.7亿参数模型在葡语任务中达当前最优表现。
- 适合研究或开发葡萄牙语AI的学者与工程师使用。
我们提出Tucano 2,一套全开源的大语言模型(LLM)系列,参数量为0.5-3.7亿,旨在填补开源葡萄牙语大模型的发展空白。继此前工作后,我们扩展并提升GigaVerbo-v2数据集的质量与规模,引入新合成数据集GigaVerbo-v2 Synth以补全数据缺口,并构建两个后训练数据集:GigaVerbo-v2 SFT与GigaVerbo-v2 Preferences,支持检索增强生成、编程、工具调用、链式推理等多领域训练。通过大量消融实验,我们设计了Tucano 2系列(Base、Instruct、Think)的预训练与持续预训练方案,在多个葡语建模基准上取得最佳性能。同时,我们改进并完善了先前提出的评估框架,形成覆盖不同训练阶段的综合性评估体系。所有相关资源——包括训练配方、日志与源码——均已开源,确保研究可复现、可访问、可拓展,助力更广泛的葡语自然语言处理社区发展。
原文摘要 · Abstract (English)
We present Tucano 2, a fully open suite of large language models (LLMs) with 0.5-3.7 billion parameters, designed to address certain gaps in open-source development for Portuguese LLMs. Following our previous works, we now extend our dataset, GigaVerbo-v2, to a new degree of quality and scale, while also introducing a new synthetic dataset, GigaVerbo-v2 Synth, aimed at filling missing gaps in GigaVerbo-v2, and two post-training datasets, GigaVerbo-v2 SFT and GigaVerbo-v2 Preferences, that allow Portuguese LLMs to be trained in domains like retrieval augmented generation, coding, tool use, chain-of-thought reasoning, and many other domains of interest. Through extensive ablation studies, we design both pretraining and continual pretraining recipes for the Tucano 2 suite (Base, Instruct, and Think), which achieve state-of-the-art performance on several Portuguese-language modeling benchmarks. We also extend and refine the evaluation harness introduced in our earlier work, yielding a comprehensive evaluation suite that provides strong signals across different pretraining, continual pretraining, and post-training regimes. All artifacts associated with Tucano 2 are openly released, including training recipes, logs, and source code, ensuring that our work is reproducible, accessible, and extendable by the broader Portuguese NLP community.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。