arXiv:2608.20387cs.CLcs.AI2026-08中稿 · Interspeech 2026

用开放指令训练语音合成,让模型更懂情绪和风格。

Poly-InstructTTS: Learning In-the-Wild Expressive Speech Synthesis from Open-Ended Instructions

论文配图:Poly-InstructTTS: Learning In-the-Wild Expressive Speech Synthesis from Open-Ended Instructions
图 1 · 摘自论文原文
  • 用多模态数据构建1000小时指令标注语料,覆盖上千种情绪风格。
  • 无需提示词的GPT架构结合流匹配模块,实现音色与表达同步控制。
  • 支持特定说话人微调,适合个性化语音助手、有声书等场景。

尽管当前文本到语音(TTS)模型已具备高自然度,但通过自然语言指令精细控制语音表现仍具挑战。我们提出Poly-InstructTTS,利用真实场景中的音视频数据,从开放式指令中学习富有表现力的语音合成。构建了一个可扩展的多模态流水线,生成了包含1,000+细粒度情感与风格的1,000小时指令标注语料库。该框架采用无提示词的GPT结构,结合基于属性的思维标记,并通过流匹配模块注入参考音频的音色特征。同时提出说话人微调方法,可在保留个人特质的前提下实现指令控制迁移。还将InstructTTSEval扩展至更广泛任务。实验表明,Poly-InstructTTS在指令遵循性和表现力方面均表现出色。音频示例与扩展测试集已在项目主页提供。

原文摘要 · Abstract (English)

While recent text-to-speech (TTS) models achieve high naturalness, controlling fine-grained expression via natural-language instructions remains challenging. We introduce Poly- InstructTTS, which learns expressive speech from open-ended instructions using in-the-wild audiovisual data. We build a scalable multi-modal pipeline to construct a 1,000-hour instruction-annotated corpus covering 1,000+ fine-grained emotions and styles. The framework uses a prompt-free GPT with attribute-based thinking tokens, followed by a flow-matching module that injects timbre from a reference audio. We also present a speaker fine-tuning procedure to transfer instruction control to specific speakers while preserving persona. We further extend InstructTTSEval with broader tasks. Experiments show that Poly-InstructTTS delivers strong performance in instruction adherence and expressiveness. Audio demos and the expanded testset are available on our project page.

语音合成指令控制多模态情绪表达

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。