PLaMo 2用高效架构和合成数据,让8B模型达到100B级日语性能。
PLaMo 2 Technical Report
- 混合Samba结构渐进转为全注意力,支持32K上下文。
- 8B模型性能媲美前代100B模型,推理效率提升显著。
- 专为日语优化,适合需要高精度日语生成的场景。
本文介绍PLaMo 2,一系列面向日语的大语言模型,采用基于Samba的混合架构,通过持续预训练逐步过渡到全注意力机制,支持32K token上下文。训练利用大量合成语料克服数据稀缺问题,计算效率通过权重复用和结构化剪枝实现。该剪枝方法使8B模型性能接近此前100B模型。后训练阶段结合监督微调(SFT)与直接偏好优化(DPO),并使用合成日语指令数据及模型融合技术进行优化。通过vLLM推理加速与量化技术,在保持最小精度损失的前提下,PLaMo 2在日语基准测试中表现领先,优于同规模开源模型,在指令遵循、语言流畅性及日语特有知识方面均具优势。
原文摘要 · Abstract (English)
In this report, we introduce PLaMo 2, a series of Japanese-focused large language models featuring a hybrid Samba-based architecture that transitions to full attention via continual pre-training to support 32K token contexts. Training leverages extensive synthetic corpora to overcome data scarcity, while computational efficiency is achieved through weight reuse and structured pruning. This efficient pruning methodology produces an 8B model that achieves performance comparable to our previous 100B model. Post-training further refines the models using a pipeline of supervised fine-tuning (SFT) and direct preference optimization (DPO), enhanced by synthetic Japanese instruction data and model merging techniques. Optimized for inference using vLLM and quantization with minimal accuracy loss, the PLaMo 2 models achieve state-of-the-art results on Japanese benchmarks, outperforming similarly-sized open models in instruction-following, language fluency, and Japanese-specific knowledge.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。