arXiv:2507.03382cs.SDeess.AS2025-07中稿 · INTERSPEECH 2025被引 2

提出跨说话人情感强度控制新方法,保持语音一致性与质量。

Speaker-agnostic Emotion Vector for Cross-speaker Emotion Intensity Control

  • 设计共享情感表达的无说话人依赖向量
  • 在未见说话人上仍能保持语音一致性与可控性
  • 适用于任意说话人的情感强度调节

跨说话人情感强度控制旨在仅使用目标说话人的中性语音,生成具有特定情感强度的语音。近期提出的感情算术方法通过单说话人情感向量实现情感强度控制,但在跨说话人场景中因源说话人与目标说话人情感向量不匹配,导致说话人一致性丢失。为此,本文提出一种无说话人依赖的情感向量,可捕捉多说话人之间的共享情感表达,适用于任意说话人。实验表明,该方法在跨说话人情感强度控制中,即使在未见说话人情况下,仍能有效保持说话人一致性、语音质量和可控性。

原文摘要 · Abstract (English)

Cross-speaker emotion intensity control aims to generate emotional speech of a target speaker with desired emotion intensities using only their neutral speech. A recently proposed method, emotion arithmetic, achieves emotion intensity control using a single-speaker emotion vector. Although this prior method has shown promising results in the same-speaker setting, it lost speaker consistency in the cross-speaker setting due to mismatches between the emotion vector of the source and target speakers. To overcome this limitation, we propose a speaker-agnostic emotion vector designed to capture shared emotional expressions across multiple speakers. This speaker-agnostic emotion vector is applicable to arbitrary speakers. Experimental results demonstrate that the proposed method succeeds in cross-speaker emotion intensity control while maintaining speaker consistency, speech quality, and controllability, even in the unseen speaker case.

情感控制语音合成跨说话人

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。