构建多通道录音回放攻击仿真框架,提升语音系统安全检测能力。
Acoustic Simulation Framework for Multi-channel Replay Speech Detection
- 用公开资源构建多通道回放攻击仿真框架
- M-ALRAD在真实数据集上实现零样本泛化性能
- 引入麦克风间相位差特征增强方向感知
回放语音攻击对语音控制系统的威胁日益严重,尤其在广泛部署语音助手的智能环境中。尽管多通道音频可提供空间线索以增强检测鲁棒性,但现有数据集和方法仍主要依赖单通道录音。以往研究指出,此类攻击在新环境中的泛化能力差,亟需生成涵盖多种声学条件的数据。为此,本文提出一个基于公开资源的声学仿真框架,用于模拟多通道回放语音配置。利用该框架,我们训练了当前最先进的多通道回放检测器 M-ALRAD,并在无需任何真实训练数据的情况下,在 ReMASC 真实录制语料库上评估其泛化能力。为更好利用空间信息,我们扩展 M-ALRAD,引入相邻麦克风对之间的互通道相位差特征,增强波束成形表示的方向性提示。合成数据集已开源:https://github.com/michaelneri/synthetic-ReMASC。
原文摘要 · Abstract (English)
Replay speech attacks pose a significant threat to voice-controlled systems, especially in smart environments where voice assistants are widely deployed. While multi-channel audio offers spatial cues that can enhance replay detection robustness, existing datasets and methods predominantly rely on single-channel recordings. Moreover, previous studies highlighted that generalization of this attack to new environments is challenging, requiring new methods for generating data encompassing various acoustic conditions. Hence, in this work we introduce an acoustic simulation framework designed to simulate multi-channel replay speech configurations using publicly available resources. Using the framework, we train the state-of-the-art multi-channel replay detector M-ALRAD and evaluate its generalisation on the ReMASC real-recording corpus without any real training data. To improve the exploitation of spatial information, we extend M-ALRAD with inter-channel phase difference features computed for adjacent microphone pairs, augmenting the beamformed representation with directional cues. Synthetic datasets are available at https://github.com/michaelneri/synthetic-ReMASC.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。