arXiv:2508.07836eess.AScs.LG2025-08中稿 · WOCCI, 2025 - Inte…

针对儿童语音验证数据少的问题,提出迭代适配器提升模型迁移效果。

G-IFT: A Gated Linear Unit adapter with Iterative Fine-Tuning for Low-Resource Children's Speaker Verification

  • 用门控线性单元适配器连接预训练模型与分类器
  • 在OGI和MyST数据集上等错误率显著降低
  • 适用于多种语音验证架构,尤其适合儿童语音场景

基于成人语音训练的说话人验证系统在儿童语音上表现不佳,主要因声学差异且儿童语音数据有限,导致微调效果差。本文提出一种新型框架G-IFT(门控线性单元适配器结合迭代微调),通过在预训练说话人嵌入模型与分类器之间插入门控线性单元适配器,并分步迭代优化分类器、适配器与预训练模型。该框架对底层语音验证架构无依赖性。在ECAPA-TDNN、ResNet和X-vector三种架构上,使用OGI和MyST数据集进行实验,结果表明G-IFT框架在所有设置下均显著降低等错误率,优于基线方法。

原文摘要 · Abstract (English)

Speaker Verification (SV) systems trained on adults speech often underperform on children's SV due to the acoustic mismatch, and limited children speech data makes fine-tuning not very effective. In this paper, we propose an innovative framework, a Gated Linear Unit adapter with Iterative Fine-Tuning (G-IFT), to enhance knowledge transfer efficiency between the high-resource adults speech domain and the low-resource children's speech domain. In this framework, a Gated Linear Unit adapter is first inserted between the pre-trained speaker embedding model and the classifier. Then the classifier, adapter, and pre-trained speaker embedding model are optimized sequentially in an iterative way. This framework is agnostic to the type of the underlying architecture of the SV system. Our experiments on ECAPA-TDNN, ResNet, and X-vector architectures using the OGI and MyST datasets demonstrate that the G-IFT framework yields consistent reductions in Equal Error Rates compared to baseline methods.

说话人验证儿童语音知识迁移适配器

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。