arXiv:2409.02239cs.SDcs.AI2024-09中稿 · IEEE SLT 2024被引 3

用时序保持的最优传输方法,提升语音识别中的跨模态知识迁移效果。

Temporal Order Preserved Optimal Transport-based Cross-modal Knowledge Transfer Learning for ASR

  • 引入时序保持的最优传输,对齐声学与语言特征的时空关系。
  • 在中文语音识别中显著优于现有基于语言模型的知识迁移方法。
  • 适合需要高精度跨模态对齐的语音识别研究者使用。

将预训练语言模型(PLM)中的语言知识迁移到声学模型中,可显著提升自动语音识别(ASR)性能。然而,由于跨模态特征分布异质性,设计有效的特征对齐与知识迁移模型仍具挑战。最优传输(OT)能有效度量概率分布差异,在跨模态对齐与知识迁移中具有潜力。但原始OT在对齐时将声学与语言特征序列视为无序集合,忽略了时间顺序信息,导致需耗时预训练以学习良好对齐。本文提出一种时序保持的最优传输(TOT)跨模态对齐与知识迁移(CAKT)框架(TOT-CAKT)。TOT-CAKT通过平滑映射声学序列的局部邻近帧到语言序列的邻近区域,保留其时间顺序关系,实现更精准的特征对齐。基于此框架,我们在使用预训练中文语言模型的普通话语音识别任务中进行了实验。结果表明,所提方法显著优于多个采用语言知识迁移的先进模型,解决了原始OT方法在序列特征对齐上的不足。

原文摘要 · Abstract (English)

Transferring linguistic knowledge from a pretrained language model (PLM) to an acoustic model has been shown to greatly improve the performance of automatic speech recognition (ASR). However, due to the heterogeneous feature distributions in cross-modalities, designing an effective model for feature alignment and knowledge transfer between linguistic and acoustic sequences remains a challenging task. Optimal transport (OT), which efficiently measures probability distribution discrepancies, holds great potential for aligning and transferring knowledge between acoustic and linguistic modalities. Nonetheless, the original OT treats acoustic and linguistic feature sequences as two unordered sets in alignment and neglects temporal order information during OT coupling estimation. Consequently, a time-consuming pretraining stage is required to learn a good alignment between the acoustic and linguistic representations. In this paper, we propose a Temporal Order Preserved OT (TOT)-based Cross-modal Alignment and Knowledge Transfer (CAKT) (TOT-CAKT) for ASR. In the TOT-CAKT, local neighboring frames of acoustic sequences are smoothly mapped to neighboring regions of linguistic sequences, preserving their temporal order relationship in feature alignment and matching. With the TOT-CAKT model framework, we conduct Mandarin ASR experiments with a pretrained Chinese PLM for linguistic knowledge transfer. Our results demonstrate that the proposed TOT-CAKT significantly improves ASR performance compared to several state-of-the-art models employing linguistic knowledge transfer, and addresses the weaknesses of the original OT-based method in sequential feature alignment for ASR.

语音识别跨模态对齐最优传输知识迁移

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。