用文本描述和参考语音实现精细可控的语音合成。
FineCombo-TTS: Collaborative and Precise Controllable Speech Synthesis Using Text Descriptions and Reference Speech

- 统一框架融合参考语音与文本描述,实现灵活控制。
- 引入条件流匹配预测器,精准建模语音变化。
- 构建结构化数据集,支持属性相对调整,适合语音设计者使用。
可控文本到语音(TTS)已成为研究重点。然而,基于参考语音或文本描述的方法缺乏灵活性和精确控制,现有联合方法耦合松散,语音建模音色而文本仅控制整体风格。我们提出 FineCombo-TTS,一个基于参考语音并由文本描述引导的统一语音合成框架,实现对声学属性的灵活与精确控制。不依赖显式属性解耦,我们学习统一的声学表征,并引入基于条件流匹配(CFM)的语音方差预测器,以文本描述为指导,建模细粒度的参考到目标语音变换。为支持相对属性控制,我们构建了 FineEdit 数据集,明确编码源到目标属性变化。实验表明,该方法实现了灵活、精确且富有表现力的可控语音合成。
原文摘要 · Abstract (English)
Controllable text-to-speech (TTS) has become a key research focus. However, methods based on either reference speech or text descriptions lack flexibility and precise control, and recent joint approaches remain loosely coupled, with speech modeling timbre and text controlling global style. We propose FineCombo-TTS, a unified framework for speech synthesis grounded in reference speech and guided by text descriptions, enabling flexible and precise control over acoustic attributes. Instead of explicit attribute disentanglement, we learn a unified acoustic representation and introduce a Conditional Flow Matching (CFM)-based Speech Variance Predictor to model fine-grained reference-to-target transformations guided by text descriptions. To support relative attribute control, we construct FineEdit, a structured paired dataset that explicitly encodes source-to-target attribute variations. Experiments demonstrate that our approach achieves flexible, precise, and expressive controllable TTS.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。