arXiv:2502.15814cs.LGcs.AI2025-02ACL被引 14

用单块学术级显卡一天训练出高质量语音语言模型

Slamming: Training a Speech Language Model on One GPU in a Day

  • 通过分析初始化与架构,结合合成数据和偏好优化提升训练效率
  • 在24小时内完成训练,性能媲美顶尖语音语言模型
  • 适合资源有限的研究者快速复现与实验

我们提出Slam,一种在单块学术级GPU上24小时内训练高质量语音语言模型(SLMs)的方法。通过实证分析模型初始化与架构、使用合成训练数据、基于合成数据的偏好优化以及调整其他所有组件实现高效训练。实验证明该方法在更多计算资源下仍具良好扩展性,以极低算力开销达到领先SLM的性能水平。研究结果远超现有计算最优预测,为语音语言模型的可行性提供乐观展望。代码、数据、模型与样本详见:https://pages.cs.huji.ac.il/adiyoss-lab/slamming。

原文摘要 · Abstract (English)

We introduce Slam, a recipe for training high-quality Speech Language Models (SLMs) on a single academic GPU in 24 hours. We do so through empirical analysis of model initialisation and architecture, synthetic training data, preference optimisation with synthetic data and tweaking all other components. We empirically demonstrate that this training recipe also scales well with more compute getting results on par with leading SLMs in a fraction of the compute cost. We hope these insights will make SLM training and research more accessible. In the context of SLM scaling laws, our results far outperform predicted compute optimal performance, giving an optimistic view to SLM feasibility. See code, data, models, samples at - https://pages.cs.huji.ac.il/adiyoss-lab/slamming .

语音模型高效训练单卡训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。