arXiv:2508.16232eess.AS2025-08被引 3

将结构化剪枝与下游微调融合,实现语音模型高效压缩。

Hybrid Pruning: In-Situ Compression of Self-Supervised Speech Models for Speaker Verification and Anti-Spoofing

  • 剪枝与微调一体化,边训练边压缩
  • 参数量减少70%且性能损失极小
  • 适合资源受限设备部署,尤其在低资源场景表现优

尽管大规模自监督学习(SSL)模型如WavLM在语音处理中取得领先性能,但其庞大的模型规模限制了在资源受限设备上的部署。现有结构化剪枝方法通常与任务特定微调分离,采用多阶段流程难以针对不同下游任务生成最优架构。本文提出一种统一框架,将结构化剪枝融入下游微调过程,实现任务性能与模型稀疏性的一体化联合优化。该方法使模型在单阶段内学习面向具体任务的压缩架构,避免复杂多阶段流程与知识蒸馏。所获剪枝模型在大规模数据集上实现最高70%参数减少,且性能几乎无损:在Vox1-O、-E、-H上分别达到0.7%、0.8%、1.6%的等错误率。此外,该方法在低资源场景下展现更强泛化能力,降低过拟合,在ASVspoof5上实现3.7%的最优等错误率。

原文摘要 · Abstract (English)

Although large-scale self-supervised learning (SSL) models like WavLM have achieved state-of-the-art performance in speech processing, their significant size impedes deployment on resource-constrained devices. While structured pruning is a key technique for model compression, existing methods typically separate it from task-specific fine-tuning. This multi-stage approach struggles to create optimal architectures tailored for diverse downstream tasks. In this work, we introduce a unified framework that integrates structured pruning into the downstream fine-tuning process. Our framework unifies these steps, jointly optimizing for task performance and model sparsity in a single stage. This allows the model to learn a compressed architecture specifically for the end task, eliminating the need for complex multi-stage pipelines and knowledge distillation. Our pruned models achieve up to a 70\% parameter reduction with negligible performance degradation on large-scale datasets, achieving equal error rates of 0.7\%, 0.8\%, and 1.6\% on Vox1-O, -E, and -H, respectively. Furthermore, our approach demonstrates improved generalization in low-resource scenarios, reducing overfitting and achieving a state-of-the-art 3.7\% EER on ASVspoof5.

模型压缩语音识别自监督学习剪枝

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。