arXiv:2608.23791eess.AScs.AI2026-08中稿 · EMNLP被引 1

让语音合成情绪自然过渡,更贴近真人情感变化。

EmoTra-TTS: Smooth Intra-Utterance Emotion Transitions for Speech Synthesis

论文配图:EmoTra-TTS: Smooth Intra-Utterance Emotion Transitions for Speech Synthesis
图 1 · 摘自论文原文
  • 分步流融合生成帧级情绪过渡音频,解决真实数据稀缺问题。
  • 用帧级情绪参数控制语音韵律,实现30%-87%的性能提升。
  • 可精确调节情绪方向与强度,适合需要细腻情感表达的应用。

心理学研究指出,人类情感是连续演化的动态过程:情绪在数秒内上升、衰减并发生转变。然而,当前情感文本到语音(TTS)系统通常对每段语音仅使用单一离散标签或静态嵌入,未能反映情感的时间连续性。尽管近期基于大语言模型(LLM)的TTS系统可能通过理解文本隐式改变语调,但这种变化既不可控,也缺乏精度以实现精准的句内情绪转换。本文提出EmoTra-TTS,解决三大挑战:(1) 采用多阶段流融合管道生成帧对齐的情绪过渡音频,克服真实句内情绪转换数据稀缺的问题;(2) 双阶段情感维度(效价-唤醒-支配度,VAD)条件机制,在LLM中指导韵律规划,并在流解码器中实现声学建模,基于帧级VAD嵌入;(3) 采用方向-幅度解耦注入结构,将情绪方向与注入强度分离,避免内容失真。EmoTra-TTS仅增加0.43%参数量,无延迟开销,情感过渡质量相对提升30%-87%,在成对偏好测试中对四款SOTA基线及两款商用系统取得64.4%-79.5%的胜率。

原文摘要 · Abstract (English)

Psychological research on emotion dynamics has established that human affect is a continuous, evolving process: emotions rise, decay, and transition within seconds. Current emotional text-to-speech (TTS) systems, however, condition on a single discrete label or static embedding per utterance, fundamentally misaligning with the temporal nature of affect. While recent LLM-based TTS systems may implicitly vary prosody through text understanding, such variation is neither explicitly controllable nor precise enough for targeted intra-utterance transitions. We address three challenges: (1) a multi-pass flow blending pipeline synthesizes frame-aligned transition audio, circumventing the scarcity of natural intra-utterance transitions; (2) dual-stage Valence-Arousal-Dominance (VAD) conditioning guides prosodic planning in the LLM and acoustic realization in the flow decoder via frame-level VAD embeddings; (3) direction-magnitude decoupled injection structurally separates emotion direction from injection magnitude, preventing content degradation. EmoTra-TTS adds only +0.43% parameters with no latency overhead, achieves 30%-87% relative improvement on emotion transition quality, corroborated by 64.4%-79.5% overall win rates in pairwise preference tests against four SOTA baselines and two commercial systems.

语音合成情绪控制流模型情感动态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。