arXiv:2601.16316eess.AScs.CL2026-01中稿 · be presented in IE…

EdgeSpot在边缘设备上实现高效高精度关键词识别,10次射击下准确率提升8.3个百分点。

EdgeSpot: Efficient and High-Performance Few-Shot Model for Keyword Spotting

  • 结合优化的声学骨干与轻量自注意力机制,提升小样本学习效率
  • 10次射击时在1%误报率下准确率达82.0%,比基线提升8.3个百分点
  • 模型仅需29.4M MACs和128k参数,适合资源受限边缘设备

我们提出一种面向边缘设备的高效少样本关键词识别模型EdgeSpot,其采用优化的基于BC-ResNet的声学主干网络,搭配可训练的逐通道能量归一化前端和轻量级时序自注意力模块。训练过程中引入自监督教师模型,并使用子中心ArcFace损失进行知识蒸馏。实验表明,EdgeSpot在固定误报率(FAR)下持续优于强基线模型。最大版本EdgeSpot-4在10次射击、1% FAR条件下,准确率从73.7%提升至82.0%,仅需29.4M MACs与128k参数。

原文摘要 · Abstract (English)

We introduce an efficient few-shot keyword spotting model for edge devices, EdgeSpot, that pairs an optimized version of a BC-ResNet-based acoustic backbone with a trainable Per-Channel Energy Normalization frontend and lightweight temporal self-attention. Knowledge distillation is utilized during training by employing a self-supervised teacher model, optimized with Sub-center ArcFace loss. This study demonstrates that the EdgeSpot model consistently provides better accuracy at a fixed false-alarm rate (FAR) than strong BC-ResNet baselines. The largest variant, EdgeSpot-4, improves the 10-shot accuracy at 1% FAR from 73.7% to 82.0%, which requires only 29.4M MACs with 128k parameters.

关键词识别边缘计算少样本学习轻量化模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。