通过分段量化保留语音情感与语调信息,提升语音合成自然度。
Segmentation-Variant Codebooks for Preservation of Paralinguistic and Prosodic Information
- 按音素、词等不同单元分段量化,生成多路离散特征流。
- 在探针任务中显著提升情感与语调信息保留率,合成语音更自然。
- 适合语音合成、情感识别等需保留韵律信息的场景。
SSL语音模型(如HuBERT)中的量化能提升语言建模、语音重合成和文本转语音任务的压缩效率与性能,但常丢失情感、重音等韵律与副语言信息。增大码本规模可缓解部分损失,但会无效增加比特率。本文提出分段变体码本(SVCs),在帧、音素、词、话语等不同语言单位上进行量化,将语音分解为多路段特定的离散特征流。实验表明,SVCs在各类探针任务中显著更有效地保留了韵律与副语言信息。此外,我们发现先聚合后量化比先量化后聚合更能保留段级信息。重合成实验进一步验证了风格表达更准确,音质略有提升,同时保持可懂度。
原文摘要 · Abstract (English)
Quantization in SSL speech models (e.g., HuBERT) improves compression and performance in tasks like language modeling, resynthesis, and text-to-speech but often discards prosodic and paralinguistic information (e.g., emotion, prominence). While increasing codebook size mitigates some loss, it inefficiently raises bitrates. We propose Segmentation-Variant Codebooks (SVCs), which quantize speech at distinct linguistic units (frame, phone, word, utterance), factorizing it into multiple streams of segment-specific discrete features. Our results show that SVCs are significantly more effective at preserving prosodic and paralinguistic information across probing tasks. Additionally, we find that pooling before rather than after discretization better retains segment-level information. Resynthesis experiments further confirm improved style realization and slightly improved quality while preserving intelligibility.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。