arXiv:2502.13473eess.ASeess.SP2025-02被引 3

用自适应波束成形检测语音重放攻击,提升复杂环境下的识别准确率。

Multi-channel Replay Speech Detection using an Adaptive Learnable Beamformer

  • 融合可学习波束成形与卷积循环网络,联合优化空间滤波与分类。
  • 在多麦克风ReMASC数据集上超越现有方法,尤其在复杂环境中提升显著。
  • 对未见过的环境有更强泛化能力,适合实际部署场景使用。

重放攻击是语音控制系统面临的严重威胁,攻击者通过播放录音语音获取非法访问权限。本文提出一种基于多通道音频特征的神经网络架构M-ALRAD,用于检测重放攻击。该方法结合可学习自适应波束成形与卷积循环神经网络,实现空间滤波与分类的联合优化。在包含四麦克风阵列配置及四种环境的ReMASC数据集上进行实验,结果表明该方法优于现有最先进水平,尤其在挑战性声学环境下表现突出。此外,相比先前研究,本方法在未见环境中的泛化能力更强。

原文摘要 · Abstract (English)

Replay attacks belong to the class of severe threats against voice-controlled systems, exploiting the easy accessibility of speech signals by recorded and replayed speech to grant unauthorized access to sensitive data. In this work, we propose a multi-channel neural network architecture called M-ALRAD for the detection of replay attacks based on spatial audio features. This approach integrates a learnable adaptive beamformer with a convolutional recurrent neural network, allowing for joint optimization of spatial filtering and classification. Experiments have been carried out on the ReMASC dataset, which is a state-of-the-art multi-channel replay speech detection dataset encompassing four microphones with diverse array configurations and four environments. Results on the ReMASC dataset show the superiority of the approach compared to the state-of-the-art and yield substantial improvements for challenging acoustic environments. In addition, we demonstrate that our approach is able to better generalize to unseen environments with respect to prior studies.

语音安全重放检测波束成形多通道

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。