arXiv:2503.07357eess.ASeess.SP2025-03中稿 · publication in EUS…被引 2

研究麦克风阵列不匹配对语音重放检测模型的影响,发现微调可显著提升跨设备泛化能力。

Impact of Microphone Array Mismatches to Learning-based Replay Speech Detection

  • 使用ReMASC数据集分析不同阵列间的性能退化,对比单/多通道配置
  • 阵列不匹配导致检测准确率下降,设备内泛化优于设备间
  • 仅需10分钟目标数据微调即可恢复性能,适合实际部署场景

本文研究基于深度神经网络的多通道重放语音检测器在不同麦克风阵列间的泛化能力。现有方法在未见阵列类型上表现较差,存在显著的训练-测试性能差异。利用ReMASC数据集,分析了设备间与设备内不匹配带来的性能退化,评估了单通道与多通道配置。此外,探索了通过微调缓解新阵列上性能损失的方法。结果表明,阵列不匹配会显著降低检测准确率,设备内泛化优于设备间;但仅需十分钟目标数据进行微调,即可有效恢复性能,为异构自动说话人验证环境中的重放检测系统部署提供重要参考。

原文摘要 · Abstract (English)

In this work, we investigate the generalization of a multi-channel learning-based replay speech detector, which employs adaptive beamforming and detection, across different microphone arrays. In general, deep neural network-based microphone array processing techniques generalize poorly to unseen array types, i.e., showing a significant training-test mismatch of performance. We employ the ReMASC dataset to analyze performance degradation due to inter- and intra-device mismatches, assessing both single- and multi-channel configurations. Furthermore, we explore fine-tuning to mitigate the performance loss when transitioning to unseen microphone arrays. Our findings reveal that array mismatches significantly decrease detection accuracy, with intra-device generalization being more robust than inter-device. However, fine-tuning with as little as ten minutes of target data can effectively recover performance, providing insights for practical deployment of replay detection systems in heterogeneous automatic speaker verification environments.

语音检测麦克风阵列模型泛化微调

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。