轻量级流水线实现嘈杂环境下的语音克隆与精准口型同步
A Lightweight Pipeline for Noisy Speech Voice Cloning and Accurate Lip Sync Synthesis
- 基于潜变量扩散模型,仅需少量样本即可实现高保真零样本语音克隆
- 采用轻量GAN架构实现实时口型同步,适配噪声和非约束场景
- 模块化设计便于扩展多模态与文本引导语音调控,适合实际部署
近年来,语音克隆与说话人头像生成技术在合成自然语音和逼真口型同步方面取得显著进展。现有方法通常依赖大规模数据集和计算密集型流程,且需纯净录音输入,难以在嘈杂或资源受限环境中应用。本文提出一种新型模块化流水线,结合Tortoise文本转语音系统——一种基于Transformer的潜变量扩散模型,可在仅提供少量样本的情况下实现高保真零样本语音克隆。同时,采用轻量级生成对抗网络架构,实现鲁棒的实时口型同步。该方案降低了对大规模预训练的依赖,支持在噪声和非约束环境下生成情感丰富的语音与口型同步,具有良好的可扩展性,适用于真实世界系统部署。
原文摘要 · Abstract (English)
Recent developments in voice cloning and talking head generation demonstrate impressive capabilities in synthesizing natural speech and realistic lip synchronization. Current methods typically require and are trained on large scale datasets and computationally intensive processes using clean studio recorded inputs that is infeasible in noisy or low resource environments. In this paper, we introduce a new modular pipeline comprising Tortoise text to speech. It is a transformer based latent diffusion model that can perform high fidelity zero shot voice cloning given only a few training samples. We use a lightweight generative adversarial network architecture for robust real time lip synchronization. The solution will contribute to many essential tasks concerning less reliance on massive pre training generation of emotionally expressive speech and lip synchronization in noisy and unconstrained scenarios. The modular structure of the pipeline allows an easy extension for future multi modal and text guided voice modulation and it could be used in real world systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。