用合成语音当训练源,真实录音当目标,实现零样本语音模仿。
MimicLM: Zero-Shot Voice Imitation through Autoregressive Modeling of Pseudo-Parallel Speech Corpora

- 用合成语音作输入,真实语音作目标,绕过数据稀缺难题。
- 在自然度上超越现有方法,语音相似度保持领先水平。
- 适合语音克隆、虚拟角色配音等需要高保真模仿的场景。
语音模仿旨在将源语音转换为匹配参考说话人音色与语调风格的同时保留语言内容。传统方法依赖三元组(源、参考、目标)数据,但此类数据极为稀缺。现有方案或设计复杂解耦架构,或借助外部系统生成伪平行训练数据,前者需精细建模,后者受限于合成语音质量。本文提出MimicLM,创新性地使用合成语音作为训练源,真实录音作为目标,使模型直接学习真实语音分布,突破合成质量瓶颈。基于此数据构造,引入交错式文本-音频建模以确保内容准确性,并通过偏好对齐后训练缓解合成数据带来的分布偏差。实验表明,MimicLM以简洁有效架构实现更优语音模仿效果,在自然度上显著优于现有方法,同时在说话人身份、口音和情感维度保持竞争力。
原文摘要 · Abstract (English)
Voice imitation aims to transform source speech to match a reference speaker's timbre and speaking style while preserving linguistic content. A straightforward approach is to train on triplets of (source, reference, target), where source and target share the same content but target matches the reference's voice characteristics, yet such data is extremely scarce. Existing approaches either employ carefully designed disentanglement architectures to bypass this data scarcity or leverage external systems to synthesize pseudo-parallel training data. However, the former requires intricate model design, and the latter faces a quality ceiling when synthetic speech is used as training targets. To address these limitations, we propose MimicLM, which takes a novel approach by using synthetic speech as training sources while retaining real recordings as targets. This design enables the model to learn directly from real speech distributions, breaking the synthetic quality ceiling. Building on this data construction approach, we incorporate interleaved text-audio modeling to guide the generation of content-accurate speech and apply post-training with preference alignment to mitigate the inherent distributional mismatch when training on synthetic data. Experiments demonstrate that MimicLM achieves superior voice imitation quality with a simple yet effective architecture, significantly outperforming existing methods in naturalness while maintaining competitive similarity scores across speaker identity, accent, and emotion dimensions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。