arXiv:2601.10629eess.AS2026-01被引 11

用自然语言设计声音,还能精细调整音调、语速等属性。

VoiceSculptor: Your Voice, Designed By You

  • 通过自然语言描述直接生成可控制的语音音色
  • 在InstructTTSEval-Zh上达到开源SOTA性能
  • 适合需要定制化语音合成的研究与开发者

尽管文本转语音(TTS)技术快速发展,开源系统仍缺乏真正指令跟随、对核心语音属性(如音调、语速、年龄、情感和风格)的细粒度控制。我们提出VoiceSculptor,一个开源统一系统,通过在单一框架中集成基于指令的声音设计与高保真语音克隆,填补这一空白。该系统可直接从自然语言描述生成可控的说话人音色,支持通过检索增强生成(RAG)进行迭代优化,并实现多维度的属性级编辑。设计后的音色被转化为提示波形,输入克隆模型以实现下游语音合成中的高保真音色迁移。VoiceSculptor在InstructTTSEval-Zh上达到开源SOTA水平,代码与预训练模型已完全开源,推动可复现的指令控制语音合成研究。

原文摘要 · Abstract (English)

Despite rapid progress in text-to-speech (TTS), open-source systems still lack truly instruction-following, fine-grained control over core speech attributes (e.g., pitch, speaking rate, age, emotion, and style). We present VoiceSculptor, an open-source unified system that bridges this gap by integrating instruction-based voice design and high-fidelity voice cloning in a single framework. It generates controllable speaker timbre directly from natural-language descriptions, supports iterative refinement via Retrieval-Augmented Generation (RAG), and provides attribute-level edits across multiple dimensions. The designed voice is then rendered into a prompt waveform and fed into a cloning model to enable high-fidelity timbre transfer for downstream speech synthesis. VoiceSculptor achieves open-source state-of-the-art (SOTA) on InstructTTSEval-Zh, and is fully open-sourced, including code and pretrained models, to advance reproducible instruction-controlled TTS research.

语音合成指令控制音色设计开源模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。