用专业训练让大模型更懂健身,准确率提升超10%。
Enhancing Fitness Intelligence through Domain-Specific LLM Post-Training

- 基于Qwen3做三阶段微调,融合健身专业知识
- 在两大认证考试中准确率最高提升12.73%
- 适合想开发垂直领域大模型的研究者
科学健身指导通常依赖人工专业人士,成本高且难普及。尽管大语言模型(LLMs)在普惠健身指导方面展现潜力,但通用模型缺乏领域知识,难以应对复杂场景。本文提出FitOne系列健身专用大模型(8B与32B参数),基于Qwen3基础模型,通过持续预训练、监督微调和强化学习三阶段后训练流程,在高质量知识工程数据集上构建。在专业健身认证考试(ACSM-EP与NSCA-CSCS)及知识推理、指令遵循等通用能力上进行全面评估。实验表明,相较于Qwen3基线模型,FitOne-8B/32B在ACSM-EP与NSCA-CSCS考试中平均提升达10.09%/9.29%和12.73%/7.01%;消融实验证明各训练阶段均必要,有效平衡了领域专精与通用能力保留。该研究推动了可信健身智能系统的发展,为构建垂直领域大模型提供范式。
原文摘要 · Abstract (English)
Scientific Fitness Coaching (SFC) is typically delivered by human professionals, making it costly and inaccessible to many. While recent advances in Large Language Models (LLMs) show considerable promise for more inclusive fitness coaching, directly deploying prevailing general-purpose LLMs in SFC reveals critical limitations. These models often lack sufficient domain-specific knowledge integration, leading to weak performance on complex SFC scenarios. In this paper, we introduce FitOne, a series of fitness LLMs (with 8B and 32B parameters) designed to improve reliability and domain specialization for SFC applications. Built upon the Qwen3 foundation models, FitOne is developed through a three-stage post-training pipeline consisting of continual pre-training, supervised fine-tuning, and reinforcement learning, using large-scale, high-quality datasets derived from rigorous knowledge engineering. We conduct comprehensive evaluations of FitOne on professional fitness certification exams, including ACSM-EP and NSCA-CSCS, as well as general capabilities such as knowledge reasoning and instruction following. Experimental results show that, while retaining strong general capabilities, FitOne-8B/32B achieves average improvements of up to 10.09%/9.29% and 12.73%/7.01% on the ACSM-EP and NSCA-CSCS exams, respectively, compared with the Qwen3 base models. Furthermore, in-depth ablation studies confirm the necessity of each training stage, highlighting the pipeline's effectiveness in balancing domain expertise enhancement with general ability retention. We believe this research advances LLM systems toward more reliable fitness intelligence and will inspire future research on developing domain-specific LLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。