通过数据优化提升语音语言模型的问答能力,效果超越更大模型。
Data-Centric Lessons To Improve Speech-Language Pretraining
- 系统性测试音频处理、合成数据构建与序列交织策略
- 3.8B参数模型在问答任务上比最大3倍的模型高10.2%准确率
- 适合关注语音预训练数据质量的研究者和工程师
语音问答(SQA)是实现有用且交互式人工智能系统的核心能力。尽管近期多个语音语言模型(SpeechLMs)被发布并专注于提升SQA性能,但缺乏对预训练数据处理与筛选的受控消融实验,难以明确影响性能的关键因素。本文通过数据驱动的探索,聚焦语音语言预训练中的三大核心问题:(1) 如何处理原始网络爬取的音频内容以用于语音-文本预训练;(2) 如何构建合成数据集以补充网络爬取数据;(3) 如何将(文本、音频)段落交织成训练序列。基于这些受控实验的发现,我们训练了一个3.8B参数的SpeechLM,名为SpeLangy,其在问答任务上的表现优于最大3倍的模型,绝对提升达10.2%。研究结果强调了有效数据整理在语音语言预训练中的关键作用,并为未来研究提供数据驱动的指导。
原文摘要 · Abstract (English)
Spoken Question-Answering (SQA) is a core capability for useful and interactive artificial intelligence systems. Recently, several speech-language models (SpeechLMs) have been released with a specific focus on improving their SQA performance. However, a lack of controlled ablations of pretraining data processing and curation makes it challenging to understand what factors account for performance, despite substantial gains from similar studies in other data modalities. In this work, we address this gap by conducting a data-centric exploration for pretraining SpeechLMs. We focus on three research questions fundamental to speech-language pretraining data: (1) how to process raw web-crawled audio content for speech-text pretraining, (2) how to construct synthetic pretraining datasets to augment web-crawled data and (3) how to interleave (text, audio) segments into training sequences. We apply the insights from our controlled data-centric ablations to pretrain a 3.8B-parameter SpeechLM, called SpeLangy, that outperforms models that are up to 3x larger by 10.2% absolute performance. We hope our findings highlight the impact of effective data curation for speech-language pretraining and guide future data-centric exploration in SpeechLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。