用6000亿体育数据训练小模型,性能媲美大模型。
OnlySportsLM: Optimizing Sports-Domain Language Models with SOTA Performance under Billion Parameters
- 用细粒度体育数据优化RWKV架构,构建196万参数小模型。
- 在体育任务上超越135M/360M模型,接近17亿/15亿大模型性能。
- 提供可复现的领域模型开发流程,适合垂直场景研究者。
本文探索仅用体育领域数据训练的小型语言模型潜力。研究发现,通过精心设计的小模型结构与海量训练数据,可突破模型规模限制。提出OnlySports系列:包含6000亿词元的OnlySports Dataset(源自FineWeb)、针对体育任务优化的RWKV架构(196M参数,20层,640维)、OnlySportsLM模型及OnlySports Benchmark。实验显示,OnlySportsLM在体育任务上相较先前135M/360M最优模型分别提升37.62%/34.08%准确率,并达到SomlLM 1.7B和Qwen 1.5B在体育领域的性能水平。该系列构建了高质量领域模型的完整工作流,为多专业领域高效AI开发提供可复制范式。
原文摘要 · Abstract (English)
This paper explores the potential of a small, domain-specific language model trained exclusively on sports-related data. We investigate whether extensive training data with specially designed small model structures can overcome model size constraints. The study introduces the OnlySports collection, comprising OnlySportsLM, OnlySports Dataset, and OnlySports Benchmark. Our approach involves: 1) creating a massive 600 billion tokens OnlySports Dataset from FineWeb, 2) optimizing the RWKV architecture for sports-related tasks, resulting in a 196M parameters model with 20-layer, 640-dimension structure, 3) training the OnlySportsLM on part of OnlySports Dataset, and 4) testing the resultant model on OnlySports Benchmark. OnlySportsLM achieves a 37.62%/34.08% accuracy improvement over previous 135M/360M state-of-the-art models and matches the performance of larger models such as SomlLM 1.7B and Qwen 1.5B in the sports domain. Additionally, the OnlySports collection presents a comprehensive workflow for building high-quality, domain-specific language models, providing a replicable blueprint for efficient AI development across various specialized fields.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。