arXiv:2508.01637eess.AS2025-08中稿 · the Interspeech 20…被引 1

提出跨年龄鲁棒的说话人验证系统,解决儿童与成人语音差异导致的识别难题。

An Age-Agnostic System for Robust Speaker Verification

  • 用域分类器分离语音中的年龄特征,构建统一的鲁棒声纹表示
  • 在OGI和VoxCeleb数据集上同时提升儿童与成人验证准确率
  • 适合需要覆盖全年龄段的实用化语音识别系统

在说话人验证(SV)中,儿童与成人的语音存在声学差异,导致以成人为训练对象的系统在儿童说话人验证(C-SV)任务上表现不佳。尽管领域自适应技术可提升C-SV性能,但常以牺牲成人说话人验证(A-SV)性能为代价。本文提出一种无年龄依赖的说话人验证(AASV)系统,在保持对成人语音高精度的同时显著提升对儿童语音的验证效果。该方法通过域分类器将年龄相关属性从语音中解耦,并利用提取的域信息扩展嵌入空间,形成跨年龄一致且具有强区分性的统一声纹表示。在OGI和VoxCeleb数据集上的实验表明,该方法有效缩小了不同年龄群体间的验证性能差距,为构建包容性强、可适配全年龄段的说话人验证系统奠定了基础。

原文摘要 · Abstract (English)

In speaker verification (SV), the acoustic mismatch between children's and adults' speech leads to suboptimal performance when adult-trained SV systems are applied to children's speaker verification (C-SV). While domain adaptation techniques can enhance performance on C-SV tasks, they often do so at the expense of significant degradation in performance on adults' SV (A-SV) tasks. In this study, we propose an Age Agnostic Speaker Verification (AASV) system that achieves robust performance across both C-SV and A-SV tasks. Our approach employs a domain classifier to disentangle age-related attributes from speech and subsequently expands the embedding space using the extracted domain information, forming a unified speaker representation that is robust and highly discriminative across age groups. Experiments on the OGI and VoxCeleb datasets demonstrate the effectiveness of our approach in bridging SV performance disparities, laying the foundation for inclusive and age-adaptive SV systems.

说话人验证跨年龄鲁棒性声纹表示

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。