通过筛选关键偏置信息,提升语音识别在不同偏置内容量下的鲁棒性。
Enhancing the Robustness of Contextual ASR to Varying Biasing Information Volumes Through Purified Semantic Correlation Joint Modeling
- 基于三级语义相关性联合建模,精准筛选最相关的偏置内容。
- 在不同长度偏置列表下,F1分数最高提升28.46%。
- 适合需要处理个性化偏置的语音识别系统开发者。
近年来,基于交叉注意力的上下文语音识别(ASR)模型在识别个性化偏置短语方面取得显著进展。然而,交叉注意力的有效性受偏置信息量变化影响,尤其当偏置列表长度显著增加时更为明显。我们发现,无论偏置列表长度如何,仅有限部分偏置信息与特定ASR中间表示最为相关。因此,通过识别并整合最相关偏置信息而非全部列表,可缓解偏置信息量变化对上下文ASR的影响。为此,我们提出纯净语义相关性联合建模(PSC-Joint)方法。该方法从粗到细定义并计算三个层次的语义相关性:列表级、短语级和词元级,并联合建模以生成交集,从而在多粒度下突出并融合最相关偏置信息。此外,为降低联合建模带来的计算开销,我们还提出基于分组竞争策略的净化机制,过滤无关偏置短语。相比基线模型,PSC-Joint在不同长度偏置列表下,AISHELL-1上平均相对F1提升达21.34%,KeSpeech上达28.46%。
原文摘要 · Abstract (English)
Recently, cross-attention-based contextual automatic speech recognition (ASR) models have made notable advancements in recognizing personalized biasing phrases. However, the effectiveness of cross-attention is affected by variations in biasing information volume, especially when the length of the biasing list increases significantly. We find that, regardless of the length of the biasing list, only a limited amount of biasing information is most relevant to a specific ASR intermediate representation. Therefore, by identifying and integrating the most relevant biasing information rather than the entire biasing list, we can alleviate the effects of variations in biasing information volume for contextual ASR. To this end, we propose a purified semantic correlation joint modeling (PSC-Joint) approach. In PSC-Joint, we define and calculate three semantic correlations between the ASR intermediate representations and biasing information from coarse to fine: list-level, phrase-level, and token-level. Then, the three correlations are jointly modeled to produce their intersection, so that the most relevant biasing information across various granularities is highlighted and integrated for contextual recognition. In addition, to reduce the computational cost introduced by the joint modeling of three semantic correlations, we also propose a purification mechanism based on a grouped-and-competitive strategy to filter out irrelevant biasing phrases. Compared with baselines, our PSC-Joint approach achieves average relative F1 score improvements of up to 21.34% on AISHELL-1 and 28.46% on KeSpeech, across biasing lists of varying lengths.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。