arXiv:2502.12900cs.CLcs.AI2025-02ACL被引 15

用1/50数据训练出更强语音-文本对齐模型,突破数据效率瓶颈。

Soundwave: Less is More for Speech-Text Alignment in LLMs

  • 设计轻量架构与高效训练策略,缩小语音与文本表征差距
  • 仅用1/50数据即超越Qwen2-Audio在语音翻译和AIR-Bench上的表现
  • 保持对话智能,适合资源受限场景下的语音大模型应用

现有端到端语音大语言模型通常依赖大规模标注数据训练,而数据高效训练尚未深入探索。本文聚焦语音与文本间的表征空间差异和序列长度不一致两大核心问题,提出Soundwave:一种高效训练策略与新颖架构相结合的方法。实验表明,Soundwave在语音翻译和AIR-Bench语音任务中均优于先进模型Qwen2-Audio,且仅使用其1/50的训练数据。进一步分析显示,Soundwave在对话任务中仍保持较强智能。项目代码已开源:https://github.com/FreedomIntelligence/Soundwave。

原文摘要 · Abstract (English)

Existing end-to-end speech large language models (LLMs) usually rely on large-scale annotated data for training, while data-efficient training has not been discussed in depth. We focus on two fundamental problems between speech and text: the representation space gap and sequence length inconsistency. We propose Soundwave, which utilizes an efficient training strategy and a novel architecture to address these issues. Results show that Soundwave outperforms the advanced Qwen2-Audio in speech translation and AIR-Bench speech tasks, using only one-fiftieth of the training data. Further analysis shows that Soundwave still retains its intelligence during conversation. The project is available at https://github.com/FreedomIntelligence/Soundwave.

语音对齐数据效率大模型训练LLM

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。