Trillion-7B用高效跨语言注意力,让韩语模型以极低成本实现强多语言能力。
Trillion 7B Technical Report
- 创新使用跨语言文档注意力机制,提升英到韩日的语言知识迁移效率。
- 仅用2万亿词元中10%的多语言数据,训练耗时59.4千小时H100 GPU(约14.8万美元)。
- 在4种语言27个基准上表现稳定,适合需要低成本高一致性多语言模型的场景。
我们提出Trillion-7B,目前最高效的韩语中心多语言大模型。其创新的跨语言文档注意力(XLDA)机制,显著提升了从英语向韩语、日语等目标语言的知识迁移效率。结合优化的数据混合策略、语言特异性过滤和定制化分词器构建,Trillion-7B在仅使用2万亿训练词元中10%的多语言数据情况下,仍保持优异性能,全量训练仅需59.4K H100 GPU小时(约14.8万美元)。在四种语言共27个基准上的全面评估表明,该模型具备稳健的多语言表现与出色的跨语言一致性。
原文摘要 · Abstract (English)
We introduce Trillion-7B, the most token-efficient Korean-centric multilingual LLM available. Our novel Cross-lingual Document Attention (XLDA) mechanism enables highly efficient and effective knowledge transfer from English to target languages like Korean and Japanese. Combined with optimized data mixtures, language-specific filtering, and tailored tokenizer construction, Trillion-7B achieves competitive performance while dedicating only 10\% of its 2T training tokens to multilingual data and requiring just 59.4K H100 GPU hours (\$148K) for full training. Comprehensive evaluations across 27 benchmarks in four languages demonstrate Trillion-7B's robust multilingual performance and exceptional cross-lingual consistency.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。