优化语音VAE的蒸馏损失,实现重建、理解与生成三者统一
On the Distillation Loss Functions of Speech VAE for Unified Reconstruction, Understanding, and Generation

- 提出联合边缘对齐的蒸馏损失,自适应加权提升多任务性能
- 在语音重建、理解、生成三项任务上均取得最优综合表现
- 适合需要统一语音表征的多任务语音系统研究者
基于变分自编码器(VAE)的连续语音表征已成为语音生成与重构的有前景替代方案,取代传统的频谱图或离散标记特征。近期研究通过与自监督学习(SSL)特征对齐,试图增强VAE潜在表示中的结构信息,以提升生成性能。然而,在考虑更多任务时,基于时间轴蒸馏的常用对齐方法是否最优仍不明确。为此,本文系统探讨了不同对齐方法,并分析其在重建、理解与生成三个维度上的影响。研究考察了蒸馏损失中的多种设计选择。大量实验表明,采用自适应加权的联合边缘对齐方法可实现最佳整体性能,同时支持可控的平衡。
原文摘要 · Abstract (English)
Continuous speech representations based on Variational Autoencoders (VAEs) have emerged as a promising alternative to traditional spectrogram or discrete token based features for speech generation and reconstruction. Recent research has tried to enrich the structural information in VAE latent representations by aligning with self-supervised learning (SSL) features, aiming for better generation performance. However, it remains unclear whether the widely-used alignment approach based on time-axis distillation is optimal when considering more tasks. To address this problem, this paper systematically explores different alignment approaches and analyzes their impact on the performances over three axes: reconstruction, understanding, and generation. We investigate various design choices in the distillation loss. Extensive experiments show that the joint-marginal alignment approach with adaptive weighting can achieve the best overall performance while allowing for a controllable balance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。