用声场图检测回放语音,轻量高效且物理可解释。
Multi-Channel Replay Speech Detection using Acoustic Maps
- 通过多通道声场图捕捉声源方向能量分布差异
- 在ReMASC数据集上仅用约6000参数达到良好检测效果
- 适用于真实场景下的语音助手反欺骗系统
回放攻击仍是自动说话人验证系统的重要漏洞,尤其在实时语音助手应用中。本文提出声场图作为多通道录音中回放语音检测的新型空间特征表示。该表示基于经典波束成形技术,在离散方位角和仰角网格上构建,编码了人声辐射与扬声器播放之间的物理差异。设计了一种轻量级卷积神经网络,以该表征为输入,在ReMASC数据集上实现竞争力表现,模型参数量约为6000个。实验表明,声场图能提供紧凑且具有物理可解释性的特征空间,适用于不同设备与声学环境下的回放攻击检测。
原文摘要 · Abstract (English)
Replay attacks remain a critical vulnerability for automatic speaker verification systems, particularly in real-time voice assistant applications. In this work, we propose acoustic maps as a novel spatial feature representation for replay speech detection from multi-channel recordings. Derived from classical beamforming over discrete azimuth and elevation grids, acoustic maps encode directional energy distributions that reflect physical differences between human speech radiation and loudspeaker-based replay. A lightweight convolutional neural network is designed to operate on this representation, achieving competitive performance on the ReMASC dataset with approximately 6k trainable parameters. Experimental results show that acoustic maps provide a compact and physically interpretable feature space for replay attack detection across different devices and acoustic environments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。