arXiv:2606.20650cs.CLcs.AI2026-06

用自然语言精准控制语音情感,支持48种情绪与强度细分。

EmoInstruct-TTS: Dual-Path Instruction-Guided Emotional Speech Synthesis

论文配图:EmoInstruct-TTS: Dual-Path Instruction-Guided Emotional Speech Synthesis
图 1 · 摘自论文原文
  • 双路径框架:先提取情绪嵌入,再通过指令生成声学感知的表达向量。
  • 在48种情绪状态上实现细粒度控制,合成语音情感更真实自然。
  • 适合需要高精度情感调控的语音合成系统研发者使用。

基于指令的可控语音合成允许用户通过自然语言指定情感。然而,现有方法多依赖粗粒度情绪标签,缺乏对细粒度强度的显式建模。我们提出 EmoInstruct-TTS,一种双路径指令引导的情感语音合成框架。引入 Emotion2embed,一个覆盖48种情感状态(含细粒度类别与强度等级)的有监督语义-声学情绪嵌入。为从自由文本指令中推断嵌入,设计了指令条件情绪流模型(ICE-Flow),生成具有声学依据的情绪表征。推断出的嵌入被整合进基于大语言模型的合成流程,实现显式情感控制的同时保持语义规划能力。实验表明,在情感可控性与语音自然度方面优于强基线。

原文摘要 · Abstract (English)

Instruction-based controllable speech synthesis enables users to specify emotions through natural language. However, existing approaches often rely on coarse emotion labels and lack explicit modeling of fine-grained intensity. We propose EmoInstruct-TTS, a dual-path instruction-guided framework for emotional speech synthesis. We introduce Emotion2embed, a supervised semantic-acoustic emotion embedding covering 48 emotional states, including fine-grained categories and intensity levels. To infer embeddings from free-form instructions, we design an Instruction-Conditioned Emotion Flow Model (ICE-Flow) that generates acoustically grounded emotion representations. The inferred embeddings are integrated into an LLM-based synthesis pipeline to provide explicit emotional control while preserving semantic planning. Experiments show improved emotional controllability and speech naturalness over strong baselines.

语音合成情感控制指令生成大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。