arXiv:2601.14012eess.AScs.AI2026-01中稿 · ICASSP 2026被引 1

MATE让音频文本嵌入具备多粒度层次,提升关键词识别精度。

MATE: Matryoshka Audio-Text Embeddings for Open-Vocabulary Keyword Spotting

  • 用嵌套前缀实现单向量多粒度嵌入,支持灵活匹配
  • 在WSJ和LibriPhrase上达当前最佳效果,推理无额外开销
  • 适合需要开放词汇关键词识别的语音交互场景

基于文本注册的开放词汇关键词检测(KWS)已成为固定短语触发的灵活替代方案。以往的逐话语匹配方法从嵌入学习角度看,仅在单一固定维度上学习嵌入。本文提出马特罗什卡音频-文本嵌入(MATE),一种双编码器框架,通过嵌套子嵌入(“前缀”)在一个向量中编码多种嵌入粒度。具体地,引入基于PCA的前缀对齐:每个前缀尺寸的完整文本嵌入经主成分分析压缩后作为教师目标,用于对齐音频与文本前缀。该对齐将关键线索集中在低维前缀,高维部分补充细节。MATE采用标准深度度量学习目标训练,对损失函数无依赖。据我们所知,这是首个将马特罗什卡风格嵌入应用于KWS的工作,在WSJ和LibriPhrase数据集上取得当前最优性能,且无推理开销。

原文摘要 · Abstract (English)

Open-vocabulary keyword spotting (KWS) with text-based enrollment has emerged as a flexible alternative to fixed-phrase triggers. Prior utterance-level matching methods, from an embedding-learning standpoint, learn embeddings at a single fixed dimensionality. We depart from this design and propose Matryoshka Audio-Text Embeddings (MATE), a dual-encoder framework that encodes multiple embedding granularities within a single vector via nested sub-embeddings ("prefixes"). Specifically, we introduce a PCA-guided prefix alignment: PCA-compressed versions of the full text embedding for each prefix size serve as teacher targets to align both audio and text prefixes. This alignment concentrates salient keyword cues in lower-dimensional prefixes, while higher dimensions add detail. MATE is trained with standard deep metric learning objectives for audio-text KWS, and is loss-agnostic. To our knowledge, this is the first application of matryoshka-style embeddings to KWS, achieving state-of-the-art results on WSJ and LibriPhrase without any inference overhead.

关键词识别多粒度嵌入音频文本匹配双编码器

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。