TASLA通过多层融合提升语音令牌的音韵保真度,适合低帧率文本对齐场景。
TASLA: Text-Aligned Speech Tokens with Multiple Layer-Aggregation
- 引入多层动态注意力,自适应融合浅层与深层语音特征
- 在约2.62赫兹下实现音韵更自然,优于旧框架TASTE
- 适用于语音合成、低码率语音压缩等需保真音韵的任务
我们提出Text-Aligned Speech Tokens with Multiple Layer-Aggregation(TASLA),一种面向低帧率与文本对齐场景的语音令牌化框架,旨在解决单一来源语音令牌在重建时丢失声学细节的问题。本文进一步揭示不同编码器层如何协作以捕捉全面的声学特征。先前的TASTE框架虽具备语言模型友好性,但难以保留声学细节。TASLA通过两个组件缓解这一权衡:多层动态注意力(MLDA),使每个文本位置能自适应融合来自冻结语音编码器的浅层与深层特征;以及有限标量量化(FSQ),一种每维独立离散化且支持平滑优化的方法。在约2.62赫兹(令牌/秒)下,TASLA在域内(LibriSpeech)和域外(EXPRESSO、Voxceleb)数据集上均持续提升音韵表现,达到与TASTE相当的高质量水平。我们进一步证明,动态层混合与频谱变化率相关,解释了为何在极强特征压缩和低帧率下,MLDA仍能保持音韵完整性。
原文摘要 · Abstract (English)
We propose Text-Aligned Speech Tokens with Multiple Layer-Aggregation (TASLA), which is a text-aligned speech tokenization framework that aims to address the problem that under a low-frame-rate and text-aligned regime, single-source speech tokens may lose acoustic details during reconstruction. On the other hand, this paper further explains how different encoder layers collaborate to capture comprehensive acoustic features for tokenization. Previous work, TASTE, proposed the text-aligned speech tokenization framework, which is a LM-friendly architecture, but struggles to capture acoustic details. We address this trade-off with two components: Multi-Layer Dynamic Attention (MLDA), which lets each text position adaptively mix shallow/deep features from a frozen speech encoder, and Finite Scalar Quantization (FSQ), a simple per-dimension discretization with smooth optimization. At about 2.62 Hz (tokens/s), TASLA consistently improves prosody and achieves competitive quality over TASTE on in-domain (LibriSpeech) and OOD (EXPRESSO, Voxceleb) sets. We further demonstrate that dynamic layer mixing is correlated with spectral flux and explains why MLDA preserves prosody under a low frame rate with extreme feature compression.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。