统一编辑语音的说话人、情绪和内容,支持从音素到词的精细修改。
UniSAE: Unified Speech Attribute Editing on Speaker, Emotion and Low-Level Content via Discrete Phonetic Posteriorgram Modelling

- 用离散音素后验图分解语音内容,实现音素级精准编辑。
- 在单框架内完成说话人、情绪与内容的联合编辑,精度达子音素级别。
- 适合需要高自由度语音定制的场景,如语音合成与个性化交互。
语音编辑旨在修改语句中特定部分,同时保留其余内容。现有方法主要聚焦于词级内容修改,通常将内容、说话人和情绪编辑视为独立任务,限制了编辑粒度与灵活性。本文提出UniSAE,一个统一的语音属性编辑框架,可在单一架构中实现从子音素到词级的可组合说话人、情绪与内容编辑。UniSAE引入离散音素后验图(DPPG)表示,将语音内容分解为编码音素身份、发音变体和持续时间的离散标记,支持直接的音素及子音素级编辑。对于更高层次的修改,采用自回归内容变换器预测编辑后的DPPG序列以实现词级内容编辑。编辑后的序列通过基于扩散的声学解码器生成语音,其条件依赖于解耦的说话人与情绪表示。实验结果表明,该统一框架支持精确的说话人与情绪控制,可在多种粒度下进行内容编辑,并在单一框架内联合修改所有三类属性。
原文摘要 · Abstract (English)
Speech editing aims to modify specific portions of an utterance while preserving the remaining speech. Existing approaches primarily focus on word-level content modification and typically treat content, speaker, and emotion editing as separate tasks, limiting both editing granularity and flexibility. We propose UniSAE, a unified speech attribute editing framework which supports composable speaker, emotion and content editing from sub-phoneme to word level within a single architecture. UniSAE introduces a Discrete Phonetic PosteriorGram (DPPG) representation that factorizes speech content into discrete tokens encoding phoneme identity, pronunciation variants, and duration, enabling direct phoneme- and sub-phoneme-level editing. For higher-level modifications, an autoregressive content transformer predicts edited DPPG sequences for word-level content editing. The edited sequences are rendered into speech by a diffusion-based acoustic decoder, conditioned on disentangled speaker and emotion representations. Experimental results demonstrate that the proposed unified framework supports precise speaker and emotion control, content editing at multiple granularities, and joint modification of all three attributes within a single framework.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。