arXiv:2509.26409eess.AS2025-09被引 2

用雷达信号实现无接触无声语音识别,准确率达91.1%。

IR-UWB Radar-Based Contactless Silent Speech Recognition with Attention-Enhanced Temporal Convolutional Networks

  • 结合时序卷积与注意力机制,直接从原始雷达信号学习特征。
  • 在50词识别任务中准确率91.1%,远超传统手工特征方法的74.0%。
  • 适合关注无接触生物信号识别与智能交互的研究者。

无声语音识别(SSR)是一种通过非声学语音相关生物信号识别语音内容的技术。本文提出一种基于注意力增强时序卷积网络的无接触红外超宽带(IR-UWB)雷达无声语音识别方法,利用深度学习直接从少量预处理的雷达信号中学习判别性表征。该架构融合时序卷积、自注意力与挤压-激励机制,有效捕捉发音动作模式。在包含50个词汇的识别任务上,采用留一会话交叉验证,本方法平均测试准确率达91.1%,显著优于传统手工特征方法的74.0%,验证了端到端学习的有效性。

原文摘要 · Abstract (English)

Silent speech recognition (SSR) is a technology that recognizes speech content from non-acoustic speech-related biosignals. This paper utilizes an attention-enhanced temporal convolutional network architecture for contactless IR-UWB radar-based SSR, leveraging deep learning to learn discriminative representations directly from minimally processed radar signals. The architecture integrates temporal convolutions with self-attention and squeeze-and-excitation mechanisms to capture articulatory patterns. Evaluated on a 50-word recognition task using leave-one-session-out cross-validation, our approach achieves an average test accuracy of 91.1\% compared to 74.0\% for the conventional hand-crafted feature method, demonstrating significant improvement through end-to-end learning.

无声语音识别雷达传感时序模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。