专为日语优化的大模型,训练数据超2万亿词,性能媲美GPT-4
PLaMo-100B: A Ground-Up Language Model Designed for Japanese Proficiency
- 从零训练1000亿参数模型,采用QK归一化和Z-Loss提升训练稳定性
- 在日语任务中表现优异,多项指标达到或接近GPT-4水平
- 适合需要高精度日语生成与理解的场景,如客服、内容创作
我们推出专为日语能力设计的大规模语言模型PLaMo-100B。该模型基于2万亿个标记从头训练,采用QK归一化和Z-Loss等技术保障训练过程稳定。通过监督微调和直接偏好优化等后训练方法进一步优化性能。基准测试显示,PLaMo-100B在日语特定任务中表现突出,部分指标可与前沿模型GPT-4相媲美。基础模型已开源,地址为https://huggingface.co/pfnet/plamo-100b。
原文摘要 · Abstract (English)
We introduce PLaMo-100B, a large-scale language model designed for Japanese proficiency. The model was trained from scratch using 2 trillion tokens, with architecture such as QK Normalization and Z-Loss to ensure training stability during the training process. Post-training techniques, including Supervised Fine-Tuning and Direct Preference Optimization, were applied to refine the model's performance. Benchmark evaluations suggest that PLaMo-100B performs well, particularly in Japanese-specific tasks, achieving results that are competitive with frontier models like GPT-4. The base model is available at https://huggingface.co/pfnet/plamo-100b.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。