arXiv:2411.07607eess.AScs.LG2024-11中稿 · ICASSP2025被引 6

用CTC压缩器实现语音文本联合训练,无需处理时长直接提升识别效果。

CJST: CTC Compressor based Joint Speech and Text Training for Decoder-Only ASR

  • 通过CTC压缩器与模态适配器实现双向语音文本对齐。
  • 在Librispeech和TED-LIUM2上达到最优性能,跨领域表现优异。
  • 适用于需要高效语音识别的场景,尤其适合无时长标注数据。

CTC压缩器是一种有效将音频编码器融入解码器仅模型的方法,已在多种语音任务中受到关注。本文提出一种基于CTC压缩器的语音与文本联合训练框架(CJST),用于解码器仅语音识别。CJST通过简单的模态适配器及CTC压缩器的多个特性——序列压缩、实时强制尖峰对齐和CTC类别嵌入——实现语音与文本的双向匹配。在Librispeech和TED-LIUM2数据集上的实验表明,CJST在不需处理时长信息的情况下实现了有效的文本注入,显著提升了域内与跨域场景下的性能。我们还对CTC压缩器进行了全面研究,涵盖不同压缩模式、边界情况处理以及在干净与噪声数据下的行为表现,揭示了适用于解码器仅模型的最鲁棒配置。

原文摘要 · Abstract (English)

CTC compressor can be an effective approach to integrate audio encoders to decoder-only models, which has gained growing interest for different speech applications. In this work, we propose a novel CTC compressor based joint speech and text training (CJST) framework for decoder-only ASR. CJST matches speech and text modalities from both directions by exploring a simple modality adaptor and several features of the CTC compressor, including sequence compression, on-the-fly forced peaky alignment and CTC class embeddings. Experimental results on the Librispeech and TED-LIUM2 corpora show that the proposed CJST achieves an effective text injection without the need of duration handling, leading to the best performance for both in-domain and cross-domain scenarios. We also provide a comprehensive study on CTC compressor, covering various compression modes, edge case handling and behavior under both clean and noisy data conditions, which reveals the most robust setting to use CTC compressor for decoder-only models.

语音识别CTC压缩联合训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。