轻量级多模态关键词检测框架,支持实时流式处理
Synaspot: A Lightweight, Streaming Multi-modal Framework for Keyword Spotting with Audio-Text Synergy
- 通过抑制语音中的说话人信息,提取与说话人无关的特征
- 融合语音与文本特征,在两个数据集上性能更优
- 仅需编码器提取特征,数学解码实现低延迟流式推理
连续语音流中的开放词汇关键词检测在诸多实际场景中具有重要价值。尽管多模态在关键词检测中的作用日益受到关注,但多模态融合带来的参数量增加以及端到端部署限制,制约了其实际应用。为此,我们提出一种轻量级、流式多模态框架。首先,聚焦多模态注册特征,减少语音注册中的说话人特异性(声纹)信息,提取说话人无关特征;其次,有效融合语音与文本特征;最后,引入仅需编码器提取特征的流式解码框架,通过三种模态表示进行数学解码。在LibriPhase和WenetPrase数据集上的实验表明,相比现有流式方法,本方法在显著更少参数量下实现更优性能。
原文摘要 · Abstract (English)
Open-vocabulary keyword spotting (KWS) in continuous speech streams holds significant practical value across a wide range of real-world applications. While increasing attention has been paid to the role of different modalities in KWS, their effectiveness has been acknowledged. However, the increased parameter cost from multimodal integration and the constraints of end-to-end deployment have limited the practical applicability of such models. To address these challenges, we propose a lightweight, streaming multi-modal framework. First, we focus on multimodal enrollment features and reduce speaker-specific (voiceprint) information in the speech enrollment to extract speaker-irrelevant characteristics. Second, we effectively fuse speech and text features. Finally, we introduce a streaming decoding framework that only requires the encoder to extract features, which are then mathematically decoded with our three modal representations. Experiments on LibriPhase and WenetPrase demonstrate the performance of our model. Compared to existing streaming approaches, our method achieves better performance with significantly fewer parameters.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。