提出统一框架,仅更新5.4%参数即可实现强鲁棒语音说话人验证。
UniPET-SPK: A Unified Framework for Parameter-Efficient Tuning of Pre-trained Speech Models for Robust Speaker Verification
- 设计动态门控融合适配器与提示词调优的统一框架。
- 在多个数据集上以5.4%参数量实现最优性能。
- 适合资源受限场景下的大模型高效微调,如司法语音分析。
预训练自监督语音模型具备优秀泛化能力,在下游任务中表现优异。然而,随着模型规模增大,传统微调因计算存储开销大且易过拟合而变得不切实际。本文研究参数高效调优(PET)方法,用于将大规模预训练自监督语音模型适配至说话人验证任务。提出三种方法:(i) 适配器调优,(ii) 提示词调优,(iii) 统一框架,该框架通过可学习门控机制动态融合前两者。首先提出Inner+Inter适配器架构,在Transformer中间层与输出嵌入处并行插入两类适配器,实现多层次特征适应。其次提出深度说话人提示法,将可训练提示令牌拼接至模型输入空间以引导适配。最后提出UniPET-SPK统一框架,融合上述两种方法,并学习不同数据集与场景下的最优调优组合。在VoxCeleb、CN-Celeb及1st 48-UTD司法数据集上的实验表明,UniPET-SPK始终优于单独使用两种方法、全量微调及其他参数高效调优方法,仅更新5.4%参数即达最优性能。
原文摘要 · Abstract (English)
With excellent generalization ability, SSL speech models have shown impressive performance on various downstream tasks in the pre-training and fine-tuning paradigm. However, as the size of pre-trained models grows, fine-tuning becomes practically unfeasible due to expanding computation and storage requirements and the risk of overfitting. This study explores parameter-efficient tuning (PET) methods for adapting large-scale pre-trained SSL speech models to speaker verification task. Correspondingly, we propose three PET methods: (i)an adapter-tuning method, (ii)a prompt-tuning method, and (iii)a unified framework that effectively incorporates adapter-tuning and prompt-tuning with a dynamically learnable gating mechanism. First, we propose the Inner+Inter Adapter framework, which inserts two types of adapters into pre-trained models, allowing for adaptation of latent features within the intermediate Transformer layers and output embeddings from all Transformer layers, through a parallel adapter design. Second, we propose the Deep Speaker Prompting method that concatenates trainable prompt tokens into the input space of pre-trained models to guide adaptation. Lastly, we propose the UniPET-SPK, a unified framework that effectively incorporates these two alternate PET methods into a single framework with a dynamic trainable gating mechanism. The proposed UniPET-SPK learns to find the optimal mixture of PET methods to match different datasets and scenarios. We conduct a comprehensive set of experiments on several datasets to validate the effectiveness of the proposed PET methods. Experimental results on VoxCeleb, CN-Celeb, and 1st 48-UTD forensic datasets demonstrate that the proposed UniPET-SPK consistently outperforms the two PET methods, fine-tuning, and other parameter-efficient tuning methods, achieving superior performance while updating only 5.4% of the parameters.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。