让语音合成可局部编辑,想改哪里改哪里。
A Unified Neural Codec Language Model for Selective Editable Text to Speech Generation
- 用编码器语言模型实现精准语音属性控制
- 在保持自然度前提下支持局部属性修改
- 适合需要灵活语音定制的场景
神经编码器语言模型通过完整模仿短语音提示的音色、语调和副语言信息,实现了出色的零样本语音合成。然而这种整体模仿限制了对单个属性的独立控制。本文提出统一的编码器语言模型 SpeechEdit,扩展零样本语音合成能力,引入选择性控制机制。默认情况下,SpeechEdit 完全复现语音提示的声学特征,但可按显式指令仅替换指定属性。为实现可控建模,SpeechEdit 在新构建的 LibriEdit 数据集上训练,该数据集源自 LibriHeavy,提供差分(差异感知)训练对。实验表明,该方法在保持自然度和鲁棒性的基础上,实现了对目标属性的灵活且局部化的控制。音频样例见 https://speech-editing.github.io/speech-editing/。
原文摘要 · Abstract (English)
Neural codec language models achieve impressive zero-shot Text-to-Speech (TTS) by fully imitating the acoustic characteristics of a short speech prompt, including timbre, prosody, and paralinguistic information. However, such holistic imitation limits their ability to isolate and control individual attributes. In this paper, we present a unified codec language model SpeechEdit that extends zero-shot TTS with a selective control mechanism. By default, SpeechEdit reproduces the complete acoustic profile inferred from the speech prompt, but it selectively overrides only the attributes specified by explicit control instructions. To enable controllable modeling, SpeechEdit is trained on our newly constructed LibriEdit dataset, which provides delta (difference-aware) training pairs derived from LibriHeavy. Experimental results show that our approach maintains naturalness and robustness while offering flexible and localized control over desired attributes. Audio samples are available at https://speech-editing.github.io/speech-editing/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。