arXiv:2608.00998eess.AS2026-08中稿 · publication at INT…

让语音情绪更贴合个人与文化差异,提升情感表达的感知准确性。

Beyond One-Size-Fits-All: Personalized and Culturally Adaptive Emotional TTS via Interactive Optimization of Individual Emotion Perception Spaces

论文配图:Beyond One-Size-Fits-All: Personalized and Culturally Adaptive Emotional TTS via Interactive Optimization of Individual Emotion Perception Spaces
图 1 · 摘自论文原文
  • 通过交互式遗传算法优化个体情绪感知空间,实现个性化情绪建模。
  • 在日、中、印尼用户测试中,个性化系统显著提升情绪感知匹配度。
  • 适合需要高情感适配性的智能对话、虚拟助手等应用场景。

对话式AI的发展推动了情感文本转语音(TTS)的研究。现有系统多依赖离散情绪标签,难以捕捉人类情感的细微差别。虽然部分模型采用如Russell唤醒-效价(A-V)模型的连续维度表示,但情感感知存在个体与文化差异,易导致模型输出与实际感知不一致。本文提出一种个性化且文化自适应的情感TTS框架,利用交互式遗传算法对个体化的A-V感知空间进行优化。该方法根据每位听众调整情绪表征,使生成语音的情绪表达更贴近其真实感知,优于使用平均A-V值的模型。在日本、中国和印度尼西亚参与者中的评估验证了个性化与文化适应的重要性,为突破‘一刀切’情感TTS提供了有效路径。

原文摘要 · Abstract (English)

The rise of conversational AI has increased interest in emotional Text-to-Speech (TTS). Most systems rely on discrete emotion labels, which fail to capture the nuanced nature of human affect. Recent models employ dimensional representations such as Russell's arousal-valence (A-V) model, offering finer control. However, emotional perception varies across individuals and cultures, which may cause mismatches between modeled and perceived emotions. We propose a personalized and culturally adaptive emotional TTS framework that performs interactive optimization of individualized A-V perception spaces using an Interactive Genetic Algorithm. By adapting emotion representations to each listener, the system produces speech with more perceptually aligned emotional expression than models using averaged A-V values. Evaluations with Japanese, Chinese, and Indonesian participants highlight the importance of personalization and cultural adaptation for moving beyond one-size-fits-all emotional TTS.

情感TTS个性化文化适应语音合成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。