用贝叶斯方法动态调整频段权重,提升语音验证在跨域场景下的稳定性。
Bayesian Learning for Domain-Invariant Speaker Verification and Anti-Spoofing
- 基于变分推断建模频段权重的不确定性,实现自适应加权
- 在跨数据集语音验证和抗合成攻击任务中性能显著优于传统方法
- 特别适合应对实际应用中音源、设备差异带来的领域漂移问题
真实场景下的自动语音验证(ASV)与抗欺骗检测性能会因领域不匹配而严重下降。松弛的逐实例频率归一化(RFN)通过时间与通道轴特征统计量对频段进行归一化,是降低声纹嵌入网络特征图领域依赖性的有效方法。本文认为不同频段应赋予不同权重,且权重受领域变化影响存在不确定性,因此提出利用变分推断建模权重后验分布,构建贝叶斯加权RFN(BWRFN)。该方法克服了固定权重RFN的局限性,在跨数据集ASV、跨文本转语音(TTS)反欺骗及鲁棒语音验证任务中均显著优于WRFN与RFN。
原文摘要 · Abstract (English)
The performance of automatic speaker verification (ASV) and anti-spoofing drops seriously under real-world domain mismatch conditions. The relaxed instance frequency-wise normalization (RFN), which normalizes the frequency components based on the feature statistics along the time and channel axes, is a promising approach to reducing the domain dependence in the feature maps of a speaker embedding network. We advocate that the different frequencies should receive different weights and that the weights' uncertainty due to domain shift should be accounted for. To these ends, we propose leveraging variational inference to model the posterior distribution of the weights, which results in Bayesian weighted RFN (BWRFN). This approach overcomes the limitations of fixed-weight RFN, making it more effective under domain mismatch conditions. Extensive experiments on cross-dataset ASV, cross-TTS anti-spoofing, and spoofing-robust ASV show that BWRFN is significantly better than WRFN and RFN.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。