arXiv:2503.02769cs.SDcs.CL2025-03ACL被引 14

用语音-文本交替预训练提升语音指令理解能力

InSerter: Speech Instruction Following with Unsupervised Interleaved Pre-training

  • 通过语音与文本随机交替预训练,让模型自然学习跨模态语义
  • 在首个语音指令基准上达到当前最优性能
  • 无需精心设计数据对,适合大规模语音模型训练

近年来,语音大语言模型(SpeechLLMs)发展迅速,但现有方法在遵循语音指令方面表现不佳,尤其在处理语音输入时智能显著下降。以往工作尝试通过表示对齐和行为对齐等方法缓解语音与文本表征间的语义不一致,但需在后训练阶段精心设计数据对。本文提出一种简单且可扩展的训练方法InSerter(Interleaved Speech-Text Representation Pre-training),通过将大规模文本语料库中随机抽取的片段经文本转语音生成语音,构建无监督的语音-文本交替序列进行预训练。模型由此学会根据给定语音段生成对应的文本延续,无需繁琐的数据设计。为系统评估语音指令跟随能力,我们引入SpeechInstructBench,首个专为语音指令任务设计的综合性基准。InSerter在该基准上取得最佳性能,并在多种语音处理任务中表现优异或具有竞争力。

原文摘要 · Abstract (English)

Recent advancements in speech large language models (SpeechLLMs) have attracted considerable attention. Nonetheless, current methods exhibit suboptimal performance in adhering to speech instructions. Notably, the intelligence of models significantly diminishes when processing speech-form input as compared to direct text-form input. Prior work has attempted to mitigate this semantic inconsistency between speech and text representations through techniques such as representation and behavior alignment, which involve the meticulous design of data pairs during the post-training phase. In this paper, we introduce a simple and scalable training method called InSerter, which stands for Interleaved Speech-Text Representation Pre-training. InSerter is designed to pre-train large-scale unsupervised speech-text sequences, where the speech is synthesized from randomly selected segments of an extensive text corpus using text-to-speech conversion. Consequently, the model acquires the ability to generate textual continuations corresponding to the provided speech segments, obviating the need for intensive data design endeavors. To systematically evaluate speech instruction-following capabilities, we introduce SpeechInstructBench, the first comprehensive benchmark specifically designed for speech-oriented instruction-following tasks. Our proposed InSerter achieves SOTA performance in SpeechInstructBench and demonstrates superior or competitive results across diverse speech processing tasks.

语音指令多模态预训练SpeechLLM

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。