arXiv:2502.10511eess.AScs.SD2025-02中稿 · ICASSP 2025被引 3

提升儿童语音验证对年龄变化的鲁棒性,解决声音随成长变音难题。

Enhancing Age-Related Robustness in Children Speaker Verification

  • 引入特征变换适配模块,融合局部与全局特征,减少过拟合。
  • 使用合成音频增强数据,提升系统对年龄变化的抗干扰能力。
  • 构建纵向语音数据集,首次实现跨年份验证性能评估。

儿童语音验证(C-SV)面临的主要挑战是儿童声音随成长发生显著变化。本文提出两种方法以增强对年龄相关变化的鲁棒性。首先,引入特征变换适配器(FTA)模块,将局部模式融入高层全局表征,降低对特定局部特征的过拟合,提升系统在跨年份验证中的表现。其次,采用合成音频增强(SAA)扩大数据多样性与规模,增强对年龄变化的鲁棒性。由于缺乏纵向语音数据集,难以评估C-SV系统的年龄鲁棒性,本文构建了一个纵向数据集,用于衡量跨年份验证性能。结合两种方法后,相较于基线,在一年、两年和三年间隔的跨年份评估集上,平均等错误率分别降低19.4%、13.0%和6.1%。

原文摘要 · Abstract (English)

One of the main challenges in children's speaker verification (C-SV) is the significant change in children's voices as they grow. In this paper, we propose two approaches to improve age-related robustness in C-SV. We first introduce a Feature Transform Adapter (FTA) module that integrates local patterns into higher-level global representations, reducing overfitting to specific local features and improving the inter-year SV performance of the system. We then employ Synthetic Audio Augmentation (SAA) to increase data diversity and size, thereby improving robustness against age-related changes. Since the lack of longitudinal speech datasets makes it difficult to measure age-related robustness of C-SV systems, we introduce a longitudinal dataset to assess inter-year verification robustness of C-SV systems. By integrating both of our proposed methods, the average equal error rate was reduced by 19.4%, 13.0%, and 6.1% in the one-year, two-year, and three-year gap inter-year evaluation sets, respectively, compared to the baseline.

语音验证儿童语音鲁棒性数据增强

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。