提升儿童语音验证对年龄变化的鲁棒性,解决声音随成长变音难题。
Enhancing Age-Related Robustness in Children Speaker Verification
- 引入特征变换适配模块,融合局部与全局特征,减少过拟合。
- 使用合成音频增强数据,提升系统对年龄变化的抗干扰能力。
- 构建纵向语音数据集,首次实现跨年份验证性能评估。
儿童语音验证(C-SV)面临的主要挑战是儿童声音随成长发生显著变化。本文提出两种方法以增强对年龄相关变化的鲁棒性。首先,引入特征变换适配器(FTA)模块,将局部模式融入高层全局表征,降低对特定局部特征的过拟合,提升系统在跨年份验证中的表现。其次,采用合成音频增强(SAA)扩大数据多样性与规模,增强对年龄变化的鲁棒性。由于缺乏纵向语音数据集,难以评估C-SV系统的年龄鲁棒性,本文构建了一个纵向数据集,用于衡量跨年份验证性能。结合两种方法后,相较于基线,在一年、两年和三年间隔的跨年份评估集上,平均等错误率分别降低19.4%、13.0%和6.1%。
原文摘要 · Abstract (English)
One of the main challenges in children's speaker verification (C-SV) is the significant change in children's voices as they grow. In this paper, we propose two approaches to improve age-related robustness in C-SV. We first introduce a Feature Transform Adapter (FTA) module that integrates local patterns into higher-level global representations, reducing overfitting to specific local features and improving the inter-year SV performance of the system. We then employ Synthetic Audio Augmentation (SAA) to increase data diversity and size, thereby improving robustness against age-related changes. Since the lack of longitudinal speech datasets makes it difficult to measure age-related robustness of C-SV systems, we introduce a longitudinal dataset to assess inter-year verification robustness of C-SV systems. By integrating both of our proposed methods, the average equal error rate was reduced by 19.4%, 13.0%, and 6.1% in the one-year, two-year, and three-year gap inter-year evaluation sets, respectively, compared to the baseline.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。