arXiv:2602.00560cs.SDeess.AS2026-02中稿 · Interspeech 2026被引 1

通过语义空间编辑语音内容,保持声音连续性。

Edit Content, Preserve Acoustics: Imperceptible Text-Based Speech Editing via Self-Consistency Rewards

  • 在语义空间中编辑文本,用流匹配解码器还原声音
  • 提升语音可懂度与感知质量,减少边界伪影
  • 适合语音编辑、口播修改等需保声的场景

不可察觉的文本驱动语音编辑通过修改文本实现语音内容更新,同时保持声音连续性。先前的声学空间方法存在内容与风格纠缠问题,导致生成不稳定和边界伪影。本文提出“编辑内容,保留声学”的框架:在稳定语义空间中进行编辑,声学实现由流匹配解码器完成。为确保感知一致性,引入自洽性奖励组相对策略优化,利用预训练文生语音模型作为隐式评判器,并结合可懂度与时长约束。实验表明,在可懂度、鲁棒性和感知质量上均优于当前最优的自回归与非自回归基线方法。

原文摘要 · Abstract (English)

Imperceptible text-based speech editing modifies spoken content through transcript manipulation while preserving acoustic continuity. Prior acoustic-space approaches suffer from content-style entanglement, causing unstable generation and boundary artifacts. We introduce a framework guided by the principle of "Edit Content, Preserve Acoustics". Editing is conducted in a stable semantic space, while acoustic realization is handled by a Flow Matching decoder. To ensure perceptual consistency, we propose Self-Consistency Rewards Group Relative Policy Optimization, which leverages a pre-trained Text-to-Speech model as an implicit critic, together with intelligibility and duration constraints. Experiments demonstrate consistent improvements over state-of-the-art autoregressive and non-autoregressive baselines in intelligibility, robustness, and perceptual quality.

语音编辑流匹配文本生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。