为智能眼镜语音安全构建新数据集并提出高效检测方法
AuthGlass: Benchmarking Voice Liveness Detection and Authentication on Smart Glasses via Comprehensive Acoustic Features
- 基于声场特征设计语音活体检测与认证模型
- 在四类任务中达当前最优性能,实测效果稳定
- 适合智能设备安全、语音防伪研究者参考
随着智能眼镜快速发展,语音交互因自然便捷被广泛采用,但易受欺骗攻击,且缺乏适用于该场景的公开语音活体检测与认证数据集。为此,我们收集了包含42名受试者、16通道音频数据的多声学模态数据集,并涵盖两类攻击样本。基于此数据,提出基于声场的活体检测方法AuthG-Live和多声学模态认证模型AuthG-Net。进一步在多种声学模态下基准测试七种活体检测方法和四种认证方法。结果表明,所提方法在四项基准任务中达到当前最优性能,消融实验验证了方法在真实环境约束下的泛化能力。最后,我们发布该数据集,命名为AuthGlass,以推动智能眼镜语音安全研究。
原文摘要 · Abstract (English)
With the rapid advancement of smart glasses, voice interaction has been widely adopted due to its naturalness and convenience. However, its practical deployment is often undermined by vulnerability to spoofing attacks, while no public dataset currently exists for voice liveness detection and authentication in smart-glasses scenarios. To address this challenge, we first collect a multi-acoustic-modal dataset comprising 16-channel audio data from 42 subjects, along with corresponding attack samples covering two attack categories. Based on insights derived from this collected data, we propose AuthG-Live, a sound-field-based voice liveness detection method, and AuthG-Net, a multi-acoustic-modal authentication model. We further benchmark seven voice liveness detection methods and four authentication methods across diverse acoustic modalities. The results demonstrate that our proposed approach achieves state-of-the-art performance on four benchmark tasks, and extensive ablation studies validate the generalizability of our methods \red{under real-world constraints}. Finally, we release this dataset, termed AuthGlass, to facilitate future research on voice liveness detection and authentication for smart glasses.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。