arXiv:2508.17031cs.SDcs.CL2025-08

根据文本动态插入语音并保持原音色,支持变长实时更新。

RephraseTTS: Dynamic Length Text based Speech Insertion with Speaker Style Transfer

  • 基于Transformer的非自回归模型,按文本长度和语速动态决定插入长度。
  • 在LibriTTS上优于现有基线,能保留说话人音色与韵律特征。
  • 适合语音编辑、实时修正场景,输出自然且连贯。

我们提出一种文本条件语音插入方法,即在输入语音中根据完整文本转录插入一段语音样本。典型应用场景是文本修改后需同步更新语音。所提方法采用基于Transformer的非自回归架构,可在推理时动态确定插入长度,依据文本内容和已有语音的语速。该方法能有效保持说话人音色、语调及频谱特性。在LibriTTS上的实验与用户研究显示,本方法优于基于现有自适应文本到语音方法的基线。此外,我们提供了大量定性结果,验证了生成语音的质量。

原文摘要 · Abstract (English)

We propose a method for the task of text-conditioned speech insertion, i.e. inserting a speech sample in an input speech sample, conditioned on the corresponding complete text transcript. An example use case of the task would be to update the speech audio when corrections are done on the corresponding text transcript. The proposed method follows a transformer-based non-autoregressive approach that allows speech insertions of variable lengths, which are dynamically determined during inference, based on the text transcript and tempo of the available partial input. It is capable of maintaining the speaker's voice characteristics, prosody and other spectral properties of the available speech input. Results from our experiments and user study on LibriTTS show that our method outperforms baselines based on an existing adaptive text to speech method. We also provide numerous qualitative results to appreciate the quality of the output from the proposed method.

语音插入风格迁移非自回归文本控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。