arXiv:2410.15620cs.SDcs.CL2024-10被引 1

多源语音数据训练下,高效融合模型并评估数据贡献。

Acoustic Model Optimization over Multiple Data Sources: Merging and Valuation

  • 分阶段训练多子集模型,用遗传与梯度优化算法融合
  • 新方法在公开数据上显著超越现有最佳模型
  • 引入沙普利值量化数据贡献,利于公平激励

由于隐私保护意识增强和语音数据规模庞大,自动语音识别(ASR)系统开发者难以再像以往那样使用完整数据训练声学模型。例如,数据可能由不同数据拥有者持有,无法共享。本文提出一种新范式解决该领域核心难题:第一阶段基于完整语音数据的不同子集分别训练多个声学模型;第二阶段采用两种新算法融合这些模型,生成高质量最终模型。首先提出遗传融合算法(GMA),虽能有效优化模型但效率较低;进一步提出基于梯度下降的优化融合算法(SOMA),显著缓解了效率瓶颈,同时保持高精度。在公共数据上的大量实验表明,所提方法显著优于当前最优水平。此外,引入沙普利值(Shapley Value)估计各模型的贡献得分,可用于评估数据价值,并为数据提供方提供合理激励。

原文摘要 · Abstract (English)

Due to the rising awareness of privacy protection and the voluminous scale of speech data, it is becoming infeasible for Automatic Speech Recognition (ASR) system developers to train the acoustic model with complete data as before. For example, the data may be owned by different curators, and it is not allowed to share with others. In this paper, we propose a novel paradigm to solve salient problems plaguing the ASR field. In the first stage, multiple acoustic models are trained based upon different subsets of the complete speech data, while in the second phase, two novel algorithms are utilized to generate a high-quality acoustic model based upon those trained on data subsets. We first propose the Genetic Merge Algorithm (GMA), which is a highly specialized algorithm for optimizing acoustic models but suffers from low efficiency. We further propose the SGD-Based Optimizational Merge Algorithm (SOMA), which effectively alleviates the efficiency bottleneck of GMA and maintains superior model accuracy. Extensive experiments on public data show that the proposed methods can significantly outperform the state-of-the-art. Furthermore, we introduce Shapley Value to estimate the contribution score of the trained models, which is useful for evaluating the effectiveness of the data and providing fair incentives to their curators.

语音识别模型融合数据估值

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。