arXiv:2410.04091cs.LGcs.SD2024-10

跨语言语音关键词检测新模型,用Transformer提升准确率和重复计数能力

Cross-Lingual Query-by-Example Spoken Term Detection: A Transformer-Based Approach

  • 基于XLSR-53和霍夫变换的跨语言语音检测框架
  • 在四种语言上比传统CNN提升19%-54%性能
  • 可精准统计目标语音中关键词出现次数,适合多语言场景

查询示例语音关键词检测(QbE-STD)通常受限于标注数据稀缺和语言特定性。本文提出一种新型无语言依赖的QbE-STD模型,结合图像处理技术与Transformer架构。通过预训练XLSR-53网络提取特征,并利用霍夫变换进行检测,该模型可在任意音频文件中搜索用户指定的语音关键词。在四种语言上的实验表明,相比基于CNN的基线模型,性能提升19%-54%。尽管处理速度优于动态时间规整(DTW),但准确性仍有提升空间。值得注意的是,该模型能准确统计目标音频中查询词的重复次数。

原文摘要 · Abstract (English)

Query-by-example spoken term detection (QbE-STD) is typically constrained by transcribed data scarcity and language specificity. This paper introduces a novel, language-agnostic QbE-STD model leveraging image processing techniques and transformer architecture. By employing a pre-trained XLSR-53 network for feature extraction and a Hough transform for detection, our model effectively searches for user-defined spoken terms within any audio file. Experimental results across four languages demonstrate significant performance gains (19-54%) over a CNN-based baseline. While processing time is improved compared to DTW, accuracy remains inferior. Notably, our model offers the advantage of accurately counting query term repetitions within the target audio.

语音检测跨语言TransformerXLSR

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。