arXiv:2409.15357eess.AScs.CL2024-09

提出新型声学建模框架,用关系思维提升语音识别准确率。

A Joint Spectro-Temporal Relational Thinking Based Acoustic Modeling Framework

  • 构建时频域关系图,捕捉语音片段间复杂关联
  • 在TIMIT数据集上实现7.82%的音素识别提升
  • 特别增强对易混淆元音的识别能力,适合语音识别研究者

关系思维指人类基于感官信号与先验知识形成内在关联认知的能力,这对理解语音至关重要,但尚未被用于任何人工智能语音识别系统。现有尝试多局限于仅在时间维度上的粗粒度话语级建模。为缩小人工系统与人类能力的差距,本文提出一种新型时频域关系思维声学建模框架。该框架首先生成大量概率图,以建模语音段在时间和频率维度上的关系;再将图中每对节点间的关联信息聚合并嵌入潜在表示,供下游任务使用。基于该框架的模型在TIMIT数据集的音素识别任务上,相较当前最先进系统取得7.82%的性能提升。深入分析表明,该方法主要提升了模型对元音的识别能力,而元音正是音素识别中最易混淆的部分。

原文摘要 · Abstract (English)

Relational thinking refers to the inherent ability of humans to form mental impressions about relations between sensory signals and prior knowledge, and subsequently incorporate them into their model of their world. Despite the crucial role relational thinking plays in human understanding of speech, it has yet to be leveraged in any artificial speech recognition systems. Recently, there have been some attempts to correct this oversight, but these have been limited to coarse utterance-level models that operate exclusively in the time domain. In an attempt to narrow the gap between artificial systems and human abilities, this paper presents a novel spectro-temporal relational thinking based acoustic modeling framework. Specifically, it first generates numerous probabilistic graphs to model the relationships among speech segments across both time and frequency domains. The relational information rooted in every pair of nodes within these graphs is then aggregated and embedded into latent representations that can be utilized by downstream tasks. Models built upon this framework outperform state-of-the-art systems with a 7.82\% improvement in phoneme recognition tasks over the TIMIT dataset. In-depth analyses further reveal that our proposed relational thinking modeling mainly improves the model's ability to recognize vowels, which are the most likely to be confused by phoneme recognizers.

语音识别关系建模时频分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。