统一音视频与跨模态对比学习,提升语音检索精度。
Joint Multimodal Contrastive Learning for Robust Spoken Term Detection and Keyword Spotting
- 联合优化音频-文本与音频-音频对比学习,共享嵌入空间
- 在词区分任务上超越现有AWE基线,支持多种语音检索任务
- 无需专用模型,适合需要高鲁棒性的语音识别场景
声学词嵌入(AWE)可提升语音检索任务如口语词检测(STD)和关键词识别(KWS)的效率。然而现有方法存在单模态监督、音视频对齐与音音对齐分离优化、需任务特化模型等问题。为此,我们提出一种联合多模态对比学习框架,在共享嵌入空间中统一声学与跨模态监督。该方法同时优化:(i) 受CLAP损失启发的音频-文本对比学习,实现音视频表征对齐;(ii) 通过深度词辨识(DWD)损失的音频-音频对比学习,增强类内紧凑性与类间分离性。所提方法在词区分任务上优于现有AWE基线,并灵活支持STD与KWS。据我们所知,这是首个此类综合性方法。
原文摘要 · Abstract (English)
Acoustic Word Embeddings (AWEs) improve the efficiency of speech retrieval tasks such as Spoken Term Detection (STD) and Keyword Spotting (KWS). However, existing approaches suffer from limitations, including unimodal supervision, disjoint optimization of audio-audio and audio-text alignment, and the need for task-specific models. To address these shortcomings, we propose a joint multimodal contrastive learning framework that unifies both acoustic and cross-modal supervision in a shared embedding space. Our approach simultaneously optimizes: (i) audio-text contrastive learning, inspired by the CLAP loss, to align audio and text representations and (ii) audio-audio contrastive learning, via Deep Word Discrimination (DWD) loss, to enhance intra-class compactness and inter-class separation. The proposed method outperforms existing AWE baselines on word discrimination task while flexibly supporting both STD and KWS. To our knowledge, this is the first comprehensive approach of its kind.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。