仅用文本数据实现语音理解对齐,提升模型泛化能力
TASU: Text-Only Alignment for Speech Understanding
- 不依赖音视频配对数据,仅用未配对文本引导跨模态对齐
- 零样本语音识别表现优异,且在MMSU上超越GLM-4-Voice等主流模型
- 适合缺乏标注音频数据的场景,尤其适用于持续学习与多任务泛化
近期语音大语言模型(Speech LLMs)的发展推动了多样化语音理解任务的统一架构。然而,现有对齐范式严重依赖大规模音视频配对数据和高计算成本训练,且在未见领域或任务上的泛化能力有限。为此,我们提出TASU(Text-only Alignment for Speech Understanding),一种仅利用未配对文本数据即可引导跨模态对齐的新范式。实验表明,TASU实现了具有竞争力的零样本语音识别性能。借助此特性,它可作为课程学习中的预训练阶段,提升语音识别的域泛化能力。最终,TASU将零样本泛化扩展至广泛语音理解任务,在MMSU基准上显著优于GLM-4-Voice与Step-Audio等主流Speech LLM,确立了其在语音大模型中高效、可扩展的对齐潜力。
原文摘要 · Abstract (English)
Recent advances in Speech Large Language Models (Speech LLMs) have paved the way for unified architectures across diverse speech understanding tasks. However, prevailing alignment paradigms rely heavily on large-scale audio-text paired data and computationally intensive training, yet often exhibit limited generalization to unseen domains or tasks. To address these limitations, we propose TASU (Text-only Alignment for Speech Understanding), a novel alignment paradigm that can leverage only unpaired text data to guide cross-modal alignment. Experiments show that TASU achieves competitive zero-shot speech recognition. Leveraging this property, it can further function as a pre-training stage in curriculum learning, enhancing domain generalization in speech recognition. Ultimately, TASU can extend its zero-shot generalization to a wide range of speech understanding tasks and notably outperforms prominent Speech LLMs including GLM-4-Voice and Step-Audio on the MMSU benchmark, establishing TASU as an efficient and scalable alignment paradigm for Speech LLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。