arXiv:2604.19330eess.AS2026-04被引 1

通过分阶段细化时间细节,让语音合成更自然且参数更少

Text-To-Speech with Chain-of-Details: modeling temporal dynamics in speech generation

论文配图:Text-To-Speech with Chain-of-Details: modeling temporal dynamics in speech generation
图 1 · 摘自论文原文
  • 分阶段逐步细化语音的时间粒度,用共享解码器高效生成
  • 在多个数据集上表现媲美主流方法,参数量显著减少
  • 无需额外音素时长预测器,底层阶段自动完成发音规划

近期文语转换(TTS)研究中,多阶段方法先预测语义标记,再生成声学标记。本文将粗到精的生成范式扩展至时间维度,提出链式细节(Chain-of-Details, CoD)框架,通过级联架构显式建模语音生成中的时间粗细动态。该方法在多阶段中逐级细化时间细节,每阶段针对特定时间粒度。所有时间细节预测均采用共享解码器,实现不同时间分辨率下的高效参数复用。值得注意的是,最低层级能自然完成音素规划,无需显式音素时长预测器。我们在多个数据集上评估并对比多种基线方法,实验表明,CoD在性能上达到竞争水平,同时参数量显著低于现有方法。结果证明,通过CoD框架显式建模时间动态可带来更自然的语音合成。

原文摘要 · Abstract (English)

Recent advances in Text-To-Speech (TTS) synthesis have seen the popularity of multi-stage approaches that first predict semantic tokens and then generate acoustic tokens. In this paper, we extend the coarse-to-fine generation paradigm to the temporal domain and introduce Chain-of-Details (CoD), a novel framework that explicitly models temporal coarse-to-fine dynamics in speech generation using a cascaded architecture. Our method progressively refines temporal details across multiple stages, with each stage targeting a specific temporal granularity. All temporal detail predictions are performed using a shared decoder, enabling efficient parameter utilization across different temporal resolutions. Notably, we observe that the lowest detail level naturally performs phonetic planning without the need for an explicit phoneme duration predictor. We evaluate our method on several datasets and compare it against several baselines. Experimental results show that CoD achieves competitive performance with significantly fewer parameters than existing approaches. Our findings demonstrate that explicit modeling of temporal dynamics with the CoD framework leads to more natural speech synthesis.

语音合成时间建模参数效率

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。