用可学习的残差令牌提升语音建模,捕捉隐含情感与音色等未标注信息。
Residual Tokens Enhance Masked Autoencoders for Speech Modeling
- 引入可训练的残差令牌,补全显式标签无法覆盖的语音细节
- 重建质量提升,保留内容与说话人相似性同时增强表现力
- 适用于语音增强,去噪时保持自然且可控
当前语音建模依赖显式属性如音高、内容和说话人身份,但这些不足以捕捉自然语音的全部丰富性。我们提出RT-MAE,一种新型掩码自编码器框架,通过引入无监督的可训练残差令牌,编码显式标签未解释的信息(如音色变化、噪声、情绪等)。实验表明,RT-MAE在重建质量上表现更优,既保持内容和说话人相似性,又增强表达力。进一步验证其在语音增强中的应用:推理时可有效去除噪声,同时保持可控性和自然度。
原文摘要 · Abstract (English)
Recent speech modeling relies on explicit attributes such as pitch, content, and speaker identity, but these alone cannot capture the full richness of natural speech. We introduce RT-MAE, a novel masked autoencoder framework that augments the supervised attributes-based modeling with unsupervised residual trainable tokens, designed to encode the information not explained by explicit labeled factors (e.g., timbre variations, noise, emotion etc). Experiments show that RT-MAE improves reconstruction quality, preserving content and speaker similarity while enhancing expressivity. We further demonstrate its applicability to speech enhancement, removing noise at inference while maintaining controllability and naturalness.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。