融合离散与自增强表示,提升多语言语音识别性能。
Fusion of Discrete Representations and Self-Augmented Representations for Multilingual Automatic Speech Recognition
- 提出新融合机制,结合两种离散表示互补信息。
- 在LibriSpeech上降低19%字符错误率,ML-SUPERB上降24%。
- 无需多模型融合,自增强技术降低推理开销。
自监督学习(SSL)模型在多种语音任务中表现优异。连续表示虽高效但计算与存储成本高;离散表示虽性能稍降,却能通过去重和子词建模降低传输与存储开销,提升序列效率。为提升离散表示在自动语音识别(ASR)中的表现,我们提出一种新型融合机制,整合两种离散表示,保留其优势并引入互补信息以增强性能。此外,我们探索了“自增强”离散表示,对单一连续SSL表示施加变换,消除对多模型融合的依赖,进一步降低推理成本。在LibriSpeech和ML-SUPERB等基准上的实验表明,相比非融合基线,字符错误率相对降低最高达19%和24%,验证了所提方法的有效性。
原文摘要 · Abstract (English)
Self-supervised learning (SSL) models have shown exceptional capabilities across various speech-processing tasks. Continuous SSL representations are effective but suffer from high computational and storage demands. On the other hand, discrete SSL representations, although with degraded performance, reduce transmission and storage costs, and improve input sequence efficiency through de-duplication and subword-modeling. To boost the performance of discrete representations for ASR, we introduce a novel fusion mechanism that integrates two discrete representations. The fusion mechanism preserves all the benefits of discrete representation while enhancing the model's performance by integrating complementary information. Additionally, we explore "self-augmented'' discrete representations, which apply transformations to a single continuous SSL representation, eliminating the fusion mechanism's dependency on multiple SSL models and further decreasing its inference costs. Experimental results on benchmarks, including LibriSpeech and ML-SUPERB, indicate up to 19% and 24% relative character error rate improvement compared with the non-fusion baseline, validating the effectiveness of our proposed methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。