arXiv:2606.22811cs.CLcs.AI2026-06

用自然语言指令生成语音,支持多角色、唱歌等多种场景。

Bagpiper-TTS: Natural Language Guided Universal Speech Synthesis

论文配图:Bagpiper-TTS: Natural Language Guided Universal Speech Synthesis
图 1 · 摘自论文原文
  • 通过自然语言推理生成包含文本和元数据的完整语音蓝图
  • 在种子语音评估集上达到1.7%的词错误率,性能媲美专用模型
  • 适合需要灵活语音生成的应用,如角色扮演、多说话人合成

传统语音合成系统依赖固定输入格式和预定义元数据字段,难以满足灵活的用户需求。本文提出Bagpiper-TTS,一种能够处理多样自然语言请求的通用语音合成系统。给定自然语言提示后,该系统首先推理用户意图,生成一个包含转写和细致元数据的丰富描述(即综合文本蓝图),再据此合成目标语音。该模型不仅支持经典语音合成任务,还可扩展至多说话人、意图转语音、角色扮演合成、歌声合成等多种应用。实验表明,Bagpiper-TTS在Seed-TTS-Eval基准上实现1.7%的词错误率,并在基于大模型评分和人工主观评价中,多项任务表现与专用模型相当。

原文摘要 · Abstract (English)

Classical TTS systems typically rely on rigid input formats and predefined metadata slots, limiting their ability to fulfill flexible user requirements. This paper introduces Bagpiper-TTS, a universal speech synthesis system that deals with diverse natural language user requests. Given a natural language prompt, Bagpiper-TTS first reasons over the users' intent to derive a rich caption, i.e., a comprehensive textual blueprint encompassing both transcription and nuanced metadata. Subsequently, this caption guides the synthesis of the target speech. Our model inherently supports a broad spectrum of tasks besides classical TTS applications, including multi-talker, intent-to-speech, role-play synthesis, singing voice synthesis, and more. Experimental results demonstrate that Bagpiper-TTS achieves an 1.7% Word Error Rate (WER) on the Seed-TTS-Eval benchmark and match the performance of dedicated models in both LLM-as-a-judge and human subjective evaluations across multiple applications.

语音合成自然语言通用模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。