让语音合成情绪自然过渡,更贴近真人情感变化。
EmoTra-TTS: Smooth Intra-Utterance Emotion Transitions for Speech Synthesis

- 分步流融合生成帧级情绪过渡音频,解决真实数据稀缺问题。
- 用帧级情绪参数控制语音韵律,实现30%-87%的性能提升。
- 可精确调节情绪方向与强度,适合需要细腻情感表达的应用。
心理学研究指出,人类情感是连续演化的动态过程:情绪在数秒内上升、衰减并发生转变。然而,当前情感文本到语音(TTS)系统通常对每段语音仅使用单一离散标签或静态嵌入,未能反映情感的时间连续性。尽管近期基于大语言模型(LLM)的TTS系统可能通过理解文本隐式改变语调,但这种变化既不可控,也缺乏精度以实现精准的句内情绪转换。本文提出EmoTra-TTS,解决三大挑战:(1) 采用多阶段流融合管道生成帧对齐的情绪过渡音频,克服真实句内情绪转换数据稀缺的问题;(2) 双阶段情感维度(效价-唤醒-支配度,VAD)条件机制,在LLM中指导韵律规划,并在流解码器中实现声学建模,基于帧级VAD嵌入;(3) 采用方向-幅度解耦注入结构,将情绪方向与注入强度分离,避免内容失真。EmoTra-TTS仅增加0.43%参数量,无延迟开销,情感过渡质量相对提升30%-87%,在成对偏好测试中对四款SOTA基线及两款商用系统取得64.4%-79.5%的胜率。
原文摘要 · Abstract (English)
Psychological research on emotion dynamics has established that human affect is a continuous, evolving process: emotions rise, decay, and transition within seconds. Current emotional text-to-speech (TTS) systems, however, condition on a single discrete label or static embedding per utterance, fundamentally misaligning with the temporal nature of affect. While recent LLM-based TTS systems may implicitly vary prosody through text understanding, such variation is neither explicitly controllable nor precise enough for targeted intra-utterance transitions. We address three challenges: (1) a multi-pass flow blending pipeline synthesizes frame-aligned transition audio, circumventing the scarcity of natural intra-utterance transitions; (2) dual-stage Valence-Arousal-Dominance (VAD) conditioning guides prosodic planning in the LLM and acoustic realization in the flow decoder via frame-level VAD embeddings; (3) direction-magnitude decoupled injection structurally separates emotion direction from injection magnitude, preventing content degradation. EmoTra-TTS adds only +0.43% parameters with no latency overhead, achieves 30%-87% relative improvement on emotion transition quality, corroborated by 64.4%-79.5% overall win rates in pairwise preference tests against four SOTA baselines and two commercial systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。