arXiv:2501.15496eess.AScs.SD2025-01中稿 · TASLP

用变分贝叶斯自适应学习,提升语音模型在不同设备和噪声下的迁移效果。

Variational Bayesian Adaptive Learning of Deep Latent Variables for Acoustic Knowledge Transfer

  • 通过优化深度模型中的潜在变量,实现跨域知识迁移。
  • 在设备和噪声适应任务中,性能优于现有方法。
  • 适用于语音识别、声景分类等需适应真实场景的应用。

本文提出一种新颖的变分贝叶斯自适应学习方法,用于解决训练与测试条件间存在的声学差异(如录音设备、环境噪声)。不同于传统贝叶斯方法对模型参数施加不确定性而引发维度灾难的问题,本方法聚焦于估计深度神经网络中数量可控的潜在变量。源域中学到的知识被编码为潜在变量的先验分布,并在贝叶斯意义下,与目标域少量适配数据结合,逼近后验分布。提出了两种后验估计策略:高斯均场变分推断与经验贝叶斯,分别应对源-目标域是否存在平行数据的情况。此外,还研究了结构关系建模以提升近似效果。在两项声学适配任务上进行了评估:1)设备适配用于声景分类;2)噪声适配用于语音指令识别。实验结果表明,所提方法在目标域数据上获得显著提升,且持续优于当前最优的知识迁移方法。

原文摘要 · Abstract (English)

In this work, we propose a novel variational Bayesian adaptive learning approach for cross-domain knowledge transfer to address acoustic mismatches between training and testing conditions, such as recording devices and environmental noise. Different from the traditional Bayesian approaches that impose uncertainties on model parameters risking the curse of dimensionality due to the huge number of parameters, we focus on estimating a manageable number of latent variables in deep neural models. Knowledge learned from a source domain is thus encoded in prior distributions of deep latent variables and optimally combined, in a Bayesian sense, with a small set of adaptation data from a target domain to approximate the corresponding posterior distributions. Two different strategies are proposed and investigated to estimate the posterior distributions: Gaussian mean-field variational inference, and empirical Bayes. These strategies address the presence or absence of parallel data in the source and target domains. Furthermore, structural relationship modeling is investigated to enhance the approximation. We evaluated our proposed approaches on two acoustic adaptation tasks: 1) device adaptation for acoustic scene classification, and 2) noise adaptation for spoken command recognition. Experimental results show that the proposed variational Bayesian adaptive learning approach can obtain good improvements on target domain data, and consistently outperforms state-of-the-art knowledge transfer methods.

知识迁移变分贝叶斯语音适应

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。