arXiv:2411.14100eess.AScs.CL2024-11中稿 · ICASSP 2025被引 6

用双向Mamba模型生成可快速检索的语音语义编码,提升口语词检测效率与泛化能力。

BEST-STD: Bidirectional Mamba-Enhanced Speech Tokenization for Spoken Term Detection

  • 采用双向Mamba架构自监督学习帧级特征,生成一致的离散语音令牌
  • 在LibriSpeech和TIMIT上超越现有基线,推理速度更快且对说话人不敏感
  • 适合需要高效处理未知词汇的口语词检测场景

口语词检测(STD)常受限于帧级特征依赖和计算量大的DTW模板匹配,影响实际应用。为此,我们提出一种新方法,将语音编码为离散、说话人无关的语义令牌,支持基于文本的快速检索,并有效处理未登录词。该方法致力于在同一词汇的不同发音中生成一致的令牌序列。我们在Mamba编码器中引入双向状态空间建模,并在自监督学习框架下训练,以学习上下文感知的帧级特征,再将其编码为离散令牌。分析表明,相比现有分词器,我们的语音令牌具有更强的说话人不变性,更适用于STD任务。在LibriSpeech和TIMIT数据集上的实证评估显示,本方法优于现有标准基线,同时具备更高效率。

原文摘要 · Abstract (English)

Spoken term detection (STD) is often hindered by reliance on frame-level features and the computationally intensive DTW-based template matching, limiting its practicality. To address these challenges, we propose a novel approach that encodes speech into discrete, speaker-agnostic semantic tokens. This facilitates fast retrieval using text-based search algorithms and effectively handles out-of-vocabulary terms. Our approach focuses on generating consistent token sequences across varying utterances of the same term. We also propose a bidirectional state space modeling within the Mamba encoder, trained in a self-supervised learning framework, to learn contextual frame-level features that are further encoded into discrete tokens. Our analysis shows that our speech tokens exhibit greater speaker invariance than those from existing tokenizers, making them more suitable for STD tasks. Empirical evaluation on LibriSpeech and TIMIT databases indicates that our method outperforms existing STD baselines while being more efficient.

语音识别语义编码Mamba口语词检测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。