arXiv:2410.00316cs.CLcs.AI2024-10EMNLP被引 23

用少量样本实现语音克隆中的精细情绪控制

EmoKnob: Enhance Voice Cloning with Fine-Grained Emotion Control

  • 基于少样本演示,通过可表达的说话人表征空间实现情绪调控
  • 支持开放文本描述的情绪控制,提升情绪表现力与自然度
  • 提出新评估指标,适合语音合成与情感计算研究者

尽管近期文语转换(TTS)技术已能生成自然且富有表现力的语音,但缺乏用户自主选择情绪及调节强度的能力。本文提出EmoKnob框架,仅需少量任意情绪的演示样本,即可在语音合成中实现精细情绪控制。该框架利用近年来基础语音克隆模型带来的丰富说话人表征空间,基于其少样本能力,提出两种方法,支持通过开放式文本描述来施加情绪控制,从而实现对多样化细微情绪的直观操控。为推动情绪语音合成领域的系统化发展,我们引入一套评估指标,用于严谨评估情绪控制框架的情感忠实性与可识别性。通过客观与主观评测,结果表明该框架能有效将情绪嵌入语音,显著优于现有商业TTS服务的情绪表现力。

原文摘要 · Abstract (English)

While recent advances in Text-to-Speech (TTS) technology produce natural and expressive speech, they lack the option for users to select emotion and control intensity. We propose EmoKnob, a framework that allows fine-grained emotion control in speech synthesis with few-shot demonstrative samples of arbitrary emotion. Our framework leverages the expressive speaker representation space made possible by recent advances in foundation voice cloning models. Based on the few-shot capability of our emotion control framework, we propose two methods to apply emotion control on emotions described by open-ended text, enabling an intuitive interface for controlling a diverse array of nuanced emotions. To facilitate a more systematic emotional speech synthesis field, we introduce a set of evaluation metrics designed to rigorously assess the faithfulness and recognizability of emotion control frameworks. Through objective and subjective evaluations, we show that our emotion control framework effectively embeds emotions into speech and surpasses emotion expressiveness of commercial TTS services.

语音合成情绪控制少样本学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。