arXiv:2509.05609cs.CLcs.LG2025-09中稿 · ICASSP 2026

提出新方法解决语音识别中声学与语言表示的不对称对齐问题。

New Insights into Optimal Alignment of Acoustic and Linguistic Representations for Knowledge Transfer in ASR

  • 将对齐视为检测任务,用非平衡最优传输实现软匹配。
  • 确保每个语言单元至少对应一个声学帧,提升识别准确率。
  • 适合需要知识迁移的语音识别系统,尤其处理噪声和冗余帧。

在自动语音识别(ASR)的知识迁移中,声学与语言表示的对齐是一个核心挑战。这种对齐具有内在结构和非对称性:多个连续声学帧通常对应一个语言标记(多对一),而某些声学过渡区域可能关联多个相邻标记(一对一)。此外,声学序列常包含无语言对应的部分,如背景噪声或静音,导致匹配不平衡。本文将对齐与匹配视为检测问题,目标是高精度、高召回地识别有意义的对应关系,确保语言标记的完整覆盖,同时灵活处理冗余或噪声声学帧。基于此,我们提出一种非平衡最优传输对齐模型,显式处理分布不匹配和结构非对称性,实现声学与语言模态间的软且部分匹配。该方法保证每个语言标记至少与一个声学观测对齐,同时允许声学到语言单位的概率性映射。我们在基于CTC的ASR系统上,结合预训练语言模型进行知识迁移实验,结果表明该方法可灵活控制匹配程度,有效提升ASR性能。

原文摘要 · Abstract (English)

Aligning acoustic and linguistic representations is a central challenge to bridge the pre-trained models in knowledge transfer for automatic speech recognition (ASR). This alignment is inherently structured and asymmetric: while multiple consecutive acoustic frames typically correspond to a single linguistic token (many-to-one), certain acoustic transition regions may relate to multiple adjacent tokens (one-to-many). Moreover, acoustic sequences often include frames with no linguistic counterpart, such as background noise or silence may lead to imbalanced matching conditions. In this work, we take a new insight to regard alignment and matching as a detection problem, where the goal is to identify meaningful correspondences with high precision and recall ensuring full coverage of linguistic tokens while flexibly handling redundant or noisy acoustic frames in transferring linguistic knowledge for ASR. Based on this new insight, we propose an unbalanced optimal transport-based alignment model that explicitly handles distributional mismatch and structural asymmetries with soft and partial matching between acoustic and linguistic modalities. Our method ensures that every linguistic token is grounded in at least one acoustic observation, while allowing for flexible, probabilistic mappings from acoustic to linguistic units. We evaluate our proposed model with experiments on an CTC-based ASR system with a pre-trained language model for knowledge transfer. Experimental results demonstrate the effectiveness of our approach in flexibly controlling degree of matching and hence to improve ASR performance.

语音识别知识迁移对齐方法

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。