arXiv:2509.24570eess.AS2025-09被引 3

构建首个指令引导的语音风格编辑数据集,支持精准可控的语音风格转换。

ISSE: An Instruction-Guided Speech Style Editing Dataset And Benchmark

  • 用大模型生成带详细指令的语音风格编辑对,提升控制精度。
  • 包含近400小时语音和超10万组源目标配对,指令多样且具体。
  • 适合语音合成、风格迁移研究者,推动自然语言控制语音生成发展。

语音风格编辑旨在修改语音的风格特征,同时保持语义内容和说话人身份不变。然而,现有方法多依赖显式标签或参考音频,限制了灵活性与可扩展性;近期基于自然语言描述的方法仍受限于指令简化和风格控制粗糙。为此,我们提出指令引导的语音风格编辑数据集(ISSE),包含近400小时语音和超过10万对源-目标语音样本,每对均配有丰富详尽的文本编辑指令。我们还构建了基于大语言模型、表达性文本转语音和语音转换技术的系统化指令驱动语音生成流程,以生成高质量配对样本。进一步地,在ISSE上训练了指令引导的自回归语音模型,并从指令遵循度、音色保留和内容一致性三方面进行评估。实验表明,相较于其他数据集,ISSE能实现更准确、可控且泛化的语音风格编辑。项目主页见:https://ychenn1.github.io/ISSE/

原文摘要 · Abstract (English)

Speech style editing refers to modifying the stylistic properties of speech while preserving its linguistic content and speaker identity. However, most existing approaches depend on explicit labels or reference audio, which limits both flexibility and scalability. More recent attempts to use natural language descriptions remain constrained by oversimplified instructions and coarse style control. To address these limitations, we introduce an Instruction-guided Speech Style Editing Dataset (ISSE). The dataset comprises nearly 400 hours of speech and over 100,000 source-target pairs, each aligned with diverse and detailed textual editing instructions. We also build a systematic instructed speech data generation pipeline leveraging large language model, expressive text-to-speech and voice conversion technologies to construct high-quality paired samples. Furthermore, we train an instruction-guided autoregressive speech model on ISSE and evaluate it in terms of instruction adherence, timbre preservation, and content consistency. Experimental results demonstrate that ISSE enables accurate, controllable, and generalizable speech style editing compared to other datasets. The project page of ISSE is available at https://ychenn1.github.io/ISSE/.

语音编辑指令控制数据集风格迁移

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。