arXiv:2601.05329cs.SDeess.AS2026-01被引 5

用微调让零样本语音合成模型实现端到端语音编辑,效果媲美大模型。

CosyEdit: Unlocking End-to-End Speech Editing Capability from Zero-Shot Text-to-Speech Models

  • 基于CosyVoice模型通过任务特化微调和互补训练,内建文本-语音对齐
  • 仅用250小时数据训练400M参数模型,性能超越数亿参数语言模型
  • 无需复杂预处理,适合快速部署的高质量语音修改场景

自动语音编辑旨在根据文本指令修改语音内容,但传统级联系统依赖显式时间对齐和复杂预处理。为此,我们提出CosyEdit,一种从CosyVoice适配而来的端到端语音编辑模型,通过任务特化后训练和互补训练范式,在保持编辑前后语音高度一致性的同时,内建文本-语音对齐机制。该模型仅在自建的GigaEdit数据集上使用250小时监督数据进行训练,拥有400M参数,表现出可靠的语音编辑性能。大量评估表明,CosyEdit不仅优于多个数十亿参数的语言模型基线,还接近现有顶尖级联系统水平。结果证明,通过后训练可从零样本语音合成模型中解锁鲁棒高效的语音编辑能力,提供一种成本低廉的端到端高质量语音编辑解决方案。代码与音频样例见https://cjy1018.github.io/CosyEditDemoPage/。

原文摘要 · Abstract (English)

Automatic speech editing aims to modify spoken content based on textual instructions, yet traditional cascade systems rely on explicit temporal alignment and complex preprocessing. To address these limitations, we propose CosyEdit, an end-to-end speech editing model adapted from CosyVoice through task-specific post-training and a complementary training paradigm, which internalizes text--speech alignment while ensuring high consistency between the speech before and after editing. Trained on only 250 hours of supervised data from our curated GigaEdit dataset, our 400M-parameter model achieves reliable speech editing performance. Extensive evaluations show that CosyEdit not only outperforms several billion-parameter language model baselines but also approaches state-of-the-art cascade systems. These results show that robust and efficient speech editing can be unlocked from a zero-shot TTS model through post-training, offering a cost-effective end-to-end solution for high-quality speech editing. Code and audio samples are available at https://cjy1018.github.io/CosyEditDemoPage/.

语音编辑端到端模型微调语音合成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。