arXiv:2505.20341eess.AScs.AI2025-05被引 1

让语音编辑保持情感一致,避免文字修改导致情绪错乱。

Towards Emotionally Consistent Text-Based Speech Editing: Introducing EmoCorrector and The ECD-TSE Dataset

  • 用检索增强生成技术提取文本情绪特征,匹配对应情绪语音
  • 在真实数据集上验证,显著提升目标情绪表达一致性
  • 适合语音合成、内容编辑等需情感连贯的场景研究者

基于文本的语音编辑(TSE)仅通过文本修改实现语音调整,无需重新录制。然而现有方法主要关注内容准确性和音质一致性,常忽略文本变化引发的情绪波动或不一致问题。为此,我们提出 EmoCorrector,一种新型后处理校正方案:通过检索增强生成(RAG)提取编辑后文本的情绪特征,检索具有相同情绪的语音样本,并合成符合预期情绪、保留说话人身份与质量的语音。为支持情绪一致性建模的训练与评估,我们构建首个面向 TSE 的情绪校正数据集 ECD-TSE,其包含多样文本变体与丰富情绪表达的《文本-语音》配对数据。主观与客观实验及全面分析表明,EmoCorrector 显著提升了目标情绪表达,有效解决当前 TSE 方法中情绪不一致的问题。代码与音频示例见 https://github.com/AI-S2-Lab/EmoCorrector。

原文摘要 · Abstract (English)

Text-based speech editing (TSE) modifies speech using only text, eliminating re-recording. However, existing TSE methods, mainly focus on the content accuracy and acoustic consistency of synthetic speech segments, and often overlook the emotional shifts or inconsistency issues introduced by text changes. To address this issue, we propose EmoCorrector, a novel post-correction scheme for TSE. EmoCorrector leverages Retrieval-Augmented Generation (RAG) by extracting the edited text's emotional features, retrieving speech samples with matching emotions, and synthesizing speech that aligns with the desired emotion while preserving the speaker's identity and quality. To support the training and evaluation of emotional consistency modeling in TSE, we pioneer the benchmarking Emotion Correction Dataset for TSE (ECD-TSE). The prominent aspect of ECD-TSE is its inclusion of $<$text, speech$>$ paired data featuring diverse text variations and a range of emotional expressions. Subjective and objective experiments and comprehensive analysis on ECD-TSE confirm that EmoCorrector significantly enhances the expression of intended emotion while addressing emotion inconsistency limitations in current TSE methods. Code and audio examples are available at https://github.com/AI-S2-Lab/EmoCorrector.

语音编辑情感一致RAG文本到语音

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。