arXiv:2409.12992cs.SDcs.AI2024-09被引 4

提升语音编辑在陌生文本下的清晰度与自然度

DiffEditor: Enhancing Speech Editing with Semantic Enrichment and Acoustic Consistency

  • 用预训练语言模型增强音素语义信息
  • 提出一阶损失函数改善编辑边界平滑性
  • 适合需要自由文本编辑的语音应用

随着基于文本的语音编辑日益普及,对无限制自由文本编辑的需求持续增长。然而,现有语音编辑技术在处理域外(OOD)文本时,面临可懂度和声学一致性显著下降的问题。本文提出 DiffEditor,一种新型语音编辑模型,通过语义增强与声学一致性建模,显著提升在 OOD 场景下的表现。为提高编辑后语音的可懂度,我们利用预训练语言模型提取的词嵌入,丰富音素嵌入的语义信息。同时,强调帧间平滑性对声学一致性的关键作用,提出一种一阶损失函数,促进编辑边界处的平滑过渡,提升整体流畅性。实验表明,该模型在域内与域外文本场景下均达到当前最优性能。

原文摘要 · Abstract (English)

As text-based speech editing becomes increasingly prevalent, the demand for unrestricted free-text editing continues to grow. However, existing speech editing techniques encounter significant challenges, particularly in maintaining intelligibility and acoustic consistency when dealing with out-of-domain (OOD) text. In this paper, we introduce, DiffEditor, a novel speech editing model designed to enhance performance in OOD text scenarios through semantic enrichment and acoustic consistency. To improve the intelligibility of the edited speech, we enrich the semantic information of phoneme embeddings by integrating word embeddings extracted from a pretrained language model. Furthermore, we emphasize that interframe smoothing properties are critical for modeling acoustic consistency, and thus we propose a first-order loss function to promote smoother transitions at editing boundaries and enhance the overall fluency of the edited speech. Experimental results demonstrate that our model achieves state-of-the-art performance in both in-domain and OOD text scenarios.

语音编辑语义增强声学一致性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。