arXiv:2501.08057eess.AScs.AI2025-01

用条件计算融合语音多视图特征,提升模型收敛速度与鲁棒性。

Optimizing Speech Multi-View Feature Fusion through Conditional Computation

  • 设计梯度敏感门控网络,动态调节不同特征的贡献。
  • 在MUSTC数据集上加速收敛,性能媲美纯谱特征模型。
  • 适合需要高效融合多源语音特征的任务场景。

近期研究表明,自监督学习(SSL)特征在多种语音任务中表现优异,提供轻量且通用的多视图语音表征。然而,我们的研究发现,尽管SSL特征能加速模型收敛,但其更新方向与传统谱特征(如FBanks)存在冲突。为此,我们提出一种基于条件计算的通用特征融合框架,包含梯度敏感门控网络和多阶段丢弃策略。该框架有效缓解特征冲突,增强模型对多视图输入的鲁棒性。通过融合SSL与谱特征,该方法在多个语音翻译任务上显著加速收敛,同时在MUSTC数据集上的性能保持与仅使用谱特征的模型相当。

原文摘要 · Abstract (English)

Recent advancements have highlighted the efficacy of self-supervised learning (SSL) features in various speech-related tasks, providing lightweight and versatile multi-view speech representations. However, our study reveals that while SSL features expedite model convergence, they conflict with traditional spectral features like FBanks in terms of update directions. In response, we propose a novel generalized feature fusion framework grounded in conditional computation, featuring a gradient-sensitive gating network and a multi-stage dropout strategy. This framework mitigates feature conflicts and bolsters model robustness to multi-view input features. By integrating SSL and spectral features, our approach accelerates convergence and maintains performance on par with spectral models across multiple speech translation tasks on the MUSTC dataset.

语音处理特征融合条件计算自监督学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。