arXiv:2604.16056cs.SDcs.AI2026-04被引 2

无需训练即可精准编辑语音,保持原声与时间一致性。

AST: Adaptive, Seamless, and Training-Free Precise Speech Editing

论文配图:AST: Adaptive, Seamless, and Training-Free Precise Speech Editing
图 1 · 摘自论文原文
  • 基于预训练模型的隐空间重组技术,无缝拼接修改段落。
  • 在未编辑区域保持高保真度,且在时间对齐上优于现有方法。
  • 适合语音克隆、内容修正等零样本场景,无需标注数据。

文本驱动的语音编辑旨在修改特定片段,同时保留说话人身份和声学上下文。现有方法通常需要昂贵的任务专属训练或微调预训练语音合成模型,但前者常导致未编辑区域保真度下降,后者则在自然度与时间对齐间存在权衡。为此,我们提出 AST 框架——一种自适应、无缝且无需训练的语音编辑方法。该框架基于预训练的文本到语音模型,利用隐空间重组技术,将保留的源段与生成的目标段拼接,确保未编辑区域的保真度。为打破质量与可控性之间的权衡,引入自适应弱事实引导(AWFG),通过调节梅尔频谱信号实现边界平滑过渡,不破坏生成流形。此外,为弥补时间对齐评估的空白,我们构建了新基准:LibriSpeech-Edit 数据集及新的词级动态时间规整(WDTW)指标。大量实验表明,AST 在内容准确率、感知质量、说话人保真度和时间对齐性上均优于现有任务专属与微调方法。值得注意的是,其在无任何任务专属训练或成对编辑数据的情况下达到当前最佳的零样本语音编辑性能,验证了隐空间重组与 AWFG 在缓解质量-可控性权衡上的有效性。

原文摘要 · Abstract (English)

Text-based speech editing aims to modify specific segments while preserving speaker identity and acoustic context. Current approaches generally involve either expensive task-specific training or adapting pre-trained Text-to-Speech (TTS) models. However, both paradigms face challenges: task-specific methods often degrade fidelity in unedited regions, whereas TTS adaptations struggle with a trade-off between editing naturalness and temporal fidelity. To address these issues, we propose AST, an Adaptive, Seamless, and Training-free speech editing framework. Built upon pre-trained TTS, AST leverages Latent Recomposition to stitch preserved source segments with synthesized targets, guaranteeing fidelity in unedited regions. To break the quality-controllability trade-off, we introduce Adaptive Weak Fact Guidance (AWFG), which modulates a mel-space signal to ensure seamless boundary transitions without disrupting the generative manifold. Furthermore, to address evaluation gaps in temporal fidelity, we propose a new benchmark suite: the LibriSpeech-Edit dataset and a novel metric, Word-level Dynamic Time Warping (WDTW). Extensive experiments demonstrate that AST consistently outperforms existing task-specific and fine-tuned speech editing methods across content accuracy, perceptual quality, speaker preservation, and temporal fidelity. Remarkably, AST achieves state-of-the-art zero-shot speech editing performance without any task-specific training or paired editing data, validating the effectiveness of latent recomposition and AWFG in bridging the quality-controllability trade-off.

语音编辑零样本生成模型时间对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。