arXiv:2506.19159cs.CLcs.SD2025-06中稿 · Interspeech2025被引 1

用文本数据联合训练语音模型,提升识别准确率且无需额外语音数据。

Enhanced Hybrid Transducer and Attention Encoder Decoder with Text Data

  • 联合训练语音与文本数据,统一多模态内部表示
  • 在Librispeech上降低5.8%~12.8%的错误率(WER)
  • 可仅用文本实现跨领域适配,适合低资源场景

提出一种联合语音与文本优化方法,用于混合变体转换器与注意力编码器-解码器(TAED)建模,以利用大量文本语料并提升语音识别准确率。联合TAED(J-TAED)同时使用语音与文本输入进行训练,推理时仅需语音数据。训练后的模型能统一不同模态的内部表示,并进一步扩展至基于文本的领域自适应。该方法有效缓解了不匹配领域任务中的数据稀缺问题,因无需语音数据即可实现。实验表明,J-TAED成功将语音与语言信息融合于单一模型,在Librispeech数据集上将词错误率(WER)降低5.8%~12.8%。在两个域外数据集(金融领域与命名实体专注领域)上评估,文本域适配分别带来15.3%和17.8%的WER下降。

原文摘要 · Abstract (English)

A joint speech and text optimization method is proposed for hybrid transducer and attention-based encoder decoder (TAED) modeling to leverage large amounts of text corpus and enhance ASR accuracy. The joint TAED (J-TAED) is trained with both speech and text input modalities together, while it only takes speech data as input during inference. The trained model can unify the internal representations from different modalities, and be further extended to text-based domain adaptation. It can effectively alleviate data scarcity for mismatch domain tasks since no speech data is required. Our experiments show J-TAED successfully integrates speech and linguistic information into one model, and reduce the WER by 5.8 ~12.8% on the Librispeech dataset. The model is also evaluated on two out-of-domain datasets: one is finance and another is named entity focused. The text-based domain adaptation brings 15.3% and 17.8% WER reduction on those two datasets respectively.

语音识别多模态融合领域自适应文本增强

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。