arXiv:2412.15712cs.CLcs.HC2024-12ACL被引 9

用对比学习训练通用语音大模型,少数据也能超越专用模型。

Contrastive Learning for Task-Independent SpeechLLM-Pretraining

  • 用对比学习对齐文本与语音表示,实现跨层统一表征。
  • 仅用10%任务数据就超过语音翻译和问答专用模型性能。
  • 适合资源有限但需快速适配新语音任务的研究者。

大型语言模型在自然语言处理中表现优异,但高效适配语音处理任务仍具挑战。直接进行特定任务微调面临过拟合、数据需求高及计算成本大的问题。为此,我们提出一种可扩展的两阶段训练方法:(1) 使用对比学习进行无任务依赖的语音预训练,实现所有层级中文本与语音表征对齐;(2) 仅需少量数据的任务特定微调。该方法优于传统语音识别预训练,在仅使用10%任务数据的情况下,性能超越专用于语音翻译和问答的模型。

原文摘要 · Abstract (English)

Large language models (LLMs) excel in natural language processing but adapting these LLMs to speech processing tasks efficiently is not straightforward. Direct task-specific fine-tuning is limited by overfitting risks, data requirements, and computational costs. To address these challenges, we propose a scalable, two-stage training approach: (1) A task-independent speech pretraining stage using contrastive learning to align text and speech representations over all layers, followed by (2) a task-specific fine-tuning stage requiring minimal data. This approach outperforms traditional ASR pretraining and enables the model to surpass models specialized on speech translation and question answering while being trained on only 10% of the task-specific data.

语音大模型对比学习预训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。