用自然语言控制语音情绪变化方向,实现零样本情感语音转换。
TRACE-EVC: Text-Guided Relative Affective Control for Zero-Shot Emotional Voice Conversion

- 通过源语音锚定的流模型预测情绪变化方向与程度
- 在多个指令下准确保持说话人身份与语音质量
- 适合需要灵活控制情绪表达的语音应用
传统情感语音转换(EVC)依赖明确的目标情绪标签或参考,仅定义目标情感状态,却忽略了转换的方向或性质。本文提出指令引导的相对情感语音转换任务,使用自然语言指令(如“让语音稍平静些”或“明显更自信”)指定以源语音为基准的情感变换,而非固定目标。为此,我们构建了包含类别转移、强度调节和开放性情感变化的TRACE-Instruct数据集。提出零样本框架TRACE-EVC,核心是Emo-Compass模块,将每次转换建模为以源语音为锚点的修正流。该方法不依赖显式目标,而是预测情感变化的方向与程度。实验表明,TRACE-EVC能准确遵循相对情绪指令,同时保留说话人身份、语言内容和语音质量,并在标准分类情感转换任务上与传统EVC系统性能相当。
原文摘要 · Abstract (English)
Traditional emotional voice conversion (EVC) conditions generation on explicit target emotions like labels or references, defining the target affective state but omitting the direction or nature of the transition. We introduce instruction-guided relative emotional voice conversion, a task where natural-language instructions specify source-conditioned affective transformations (e.g., "make the speech slightly calmer" or "sound noticeably more confident") instead of fixed targets. To support this task, we construct TRACE-Instruct, a dataset of relative emotion instructions covering categorical transitions, intensity modifications, and open-ended affective changes. We propose TRACE-EVC, a zero-shot framework built around Emo-Compass, a module that models each conversion as a source-anchored rectified flow. Rather than conditioning on an explicit target, it predicts the direction and degree of the affective change. Experiments demonstrate that TRACE-EVC accurately follows relative emotion instructions while preserving speaker identity, linguistic content, and speech quality, and remains competitive with conventional EVC systems on standard categorical emotion conversion.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。