arXiv:2504.02398cs.CLcs.SD2025-04中稿 · COLM被引 12

interleaved 语音模型比传统模型更省算力,适合高效训练

Scaling Analysis of Interleaved Speech-Text Language Models

  • 用语音文本交替预训练,实现知识迁移
  • 算力增加时模型性能提升更快,可少用数据
  • 适合追求高效训练的语音模型研究者

现有语音语言模型(SLM)的缩放分析显示其需要远超文本模型的算力和数据,引发可行性质疑。但现代SLM常通过语音-文本交错预训练从文本语言模型(TextLM)初始化,实现知识迁移。本文通过训练数十个交错式SLM并分析缩放趋势,发现此类模型在算力增加时效率更高。结果表明,其缩放规律与无文本模型显著不同,应将更多算力预算分配给扩大模型规模而非增加训练样本。我们还研究了合成数据与TextLM架构的作用,发现所训练模型在语义语音表现上媲美领先模型,且使用更少算力和数据。模型、样例与数据已开源。

原文摘要 · Abstract (English)

Existing Speech Language Model (SLM) scaling analysis paints a bleak picture. It predicts that SLMs require much more compute and data compared to text, leading some to question the feasibility of training high-quality SLMs. However, modern SLMs are often initialised from pre-trained TextLMs using speech-text interleaving to allow knowledge transfer. This raises the question - "Do interleaved SLMs scale more efficiently than textless-SLMs?" In this paper we answer a resounding yes! We conduct scaling analysis of interleaved SLMs by training several dozen and analysing the scaling trends. We see that under this setup SLMs scale more efficiently with compute. Additionally, our results indicate that the scaling dynamics significantly differ from textless-SLMs, suggesting one should allocate notably more of the compute budget to increasing model size over training tokens. We also study the role of synthetic data and TextLM model families in unlocking this potential. Results suggest that our scaled up model achieves comparable semantic speech performance to leading models, while using less compute and data. We open source models, samples, and data - https://pages.cs.huji.ac.il/adiyoss-lab/sims/ .

语音建模模型缩放知识迁移

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。