arXiv:2508.02038cs.CLcs.SD2025-08被引 4

统一框架实现语音克隆与情绪控制,生成自然且情感丰富的中文语音。

Marco-Voice Technical Report

  • 通过对比学习分离说话人与情绪特征,独立控制语音身份与情感。
  • 在10小时中文情感语料上训练,主观评测表现优于现有方法。
  • 适合需要高保真语音合成与情绪调节的应用场景。

本文提出一个多任务语音合成系统Marco-Voice,将语音克隆与情绪可控语音合成整合于统一框架中,旨在解决跨语言与情绪情境下生成高表达性、可控制且自然的语音这一长期挑战。提出一种基于批次内对比学习的说话人-情绪解耦机制,实现说话人身份与情感风格的独立操控,并引入旋转式情感嵌入方法以实现平滑的情绪调节。为支持全面训练与评估,构建了高质量中文情感语音数据集CSEMOTIONS,包含10小时来自10位专业主播的7类情绪语音。大量实验表明,Marco-Voice在客观与主观指标上均取得显著提升,语音清晰度与情感丰富度表现优异,代表了表达性神经语音合成领域的重要进展。代码与数据集已公开于https://github.com/AIDC-AI/Marco-Voice 和 https://huggingface.co/datasets/AIDC-AI/CSEMOTIONS。

原文摘要 · Abstract (English)

This paper presents a multifunctional speech synthesis system that integrates voice cloning and emotion control speech synthesis within a unified framework. The goal of this work is to address longstanding challenges in achieving highly expressive, controllable, and natural speech generation that faithfully preserves speaker identity across diverse linguistic and emotional contexts. Our approach introduces an effective speaker-emotion disentanglement mechanism with in-batch contrastive learning, enabling independent manipulation of speaker identity and eemotional style, as well as rotational emotional embedding integration method for smooth emotion control. To support comprehensive training and evaluation, we construct CSEMOTIONS, a high-quality emotional speech dataset containing 10 hours of Mandarin speech from ten professional speakers across seven emotional categories. Extensive experiments demonstrate that our system, Marco-Voice, achieves substantial improvements in both objective and subjective metrics. Comprehensive evaluations and analysis were conducted, results show that MarcoVoice delivers competitive performance in terms of speech clarity and emotional richness, representing a substantial advance in the field of expressive neural speech synthesis. Our code and dataset are publicly available at https://github.com/AIDC-AI/Marco-Voice and https://huggingface.co/datasets/AIDC-AI/CSEMOTIONS respectively.

语音合成情绪控制语音克隆中文语音

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。