让语音和文字在共享空间中对齐,提升开放词汇关键词检测效果
Adversarial Deep Metric Learning for Cross-Modal Audio-Text Alignment in Open-Vocabulary Keyword Spotting
- 用对抗学习缩小语音与文本的表示差异,实现跨模态对齐
- 在WSJ和LibriPhrase数据集上,关键词检测准确率显著提升
- 适合做多模态语音识别、开放词汇检索的研究者参考
针对基于文本注册的开放词汇关键词检测(KWS),传统方法通常在音素或语句层面比较声学与文本嵌入。为此,我们采用深度度量学习(DML)优化声学与文本编码器,使多模态嵌入可在共享空间中直接比较。然而,语音与文本模态间的固有异质性带来显著挑战。为此,我们提出模态对抗学习(MAL),通过对抗训练模态分类器,促使两个编码器生成模态不变的嵌入表示。此外,我们应用DML实现音素级的声文对齐,并在多个DML目标间进行广泛对比。在华尔街日报(WSJ)和LibriPhrase数据集上的实验表明,该方法有效提升了性能。
原文摘要 · Abstract (English)
For text enrollment-based open-vocabulary keyword spotting (KWS), acoustic and text embeddings are typically compared at either the phoneme or utterance level. To facilitate this, we optimize acoustic and text encoders using deep metric learning (DML), enabling direct comparison of multi-modal embeddings in a shared embedding space. However, the inherent heterogeneity between audio and text modalities presents a significant challenge. To address this, we propose Modality Adversarial Learning (MAL), which reduces the domain gap in heterogeneous modality representations. Specifically, we train a modality classifier adversarially to encourage both encoders to generate modality-invariant embeddings. Additionally, we apply DML to achieve phoneme-level alignment between audio and text, and conduct extensive comparisons across various DML objectives. Experiments on the Wall Street Journal (WSJ) and LibriPhrase datasets demonstrate the effectiveness of the proposed approach.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。