arXiv:2501.15417cs.SDcs.AI2025-01中稿 · IEEE TASLP 2025被引 46

统一语音增强模型,支持多种任务且无需微调。

AnyEnhance: A Unified Generative Model with Prompt-Guidance and Self-Critic for Voice Enhancement

  • 基于掩码生成模型,通过提示引导实现上下文学习。
  • 在去噪、混响消除等任务上优于现有方法,主观听感更佳。
  • 适合需要多任务语音增强的开发者和研究人员。

我们提出 AnyEnhance,一种统一的生成式语音增强模型,可处理语音与歌唱声音。基于掩码生成模型,AnyEnhance 能同时执行去噪、去混响、去削波、超分辨率及目标说话人提取等多种任务,无需微调。引入提示引导机制,支持以参考音频的音色作为输入,提升性能并原生实现目标说话人提取。此外,通过自批评机制在生成过程中进行迭代自我评估与优化,显著提升输出质量。大量实验表明,AnyEnhance 在客观指标与主观听感测试中均优于现有方法。演示音频已公开于 https://amphionspace.github.io/anyenhance,开源代码见 https://github.com/viewfinder-annn/anyenhance-v1-ccf-aatc。

原文摘要 · Abstract (English)

We introduce AnyEnhance, a unified generative model for voice enhancement that processes both speech and singing voices. Based on a masked generative model, AnyEnhance is capable of handling both speech and singing voices, supporting a wide range of enhancement tasks including denoising, dereverberation, declipping, super-resolution, and target speaker extraction, all simultaneously and without fine-tuning. AnyEnhance introduces a prompt-guidance mechanism for in-context learning, which allows the model to natively accept a reference speaker's timbre. In this way, it could boost enhancement performance when a reference audio is available and enable the target speaker extraction task without altering the underlying architecture. Moreover, we also introduce a self-critic mechanism into the generative process for masked generative models, yielding higher-quality outputs through iterative self-assessment and refinement. Extensive experiments on various enhancement tasks demonstrate AnyEnhance outperforms existing methods in terms of both objective metrics and subjective listening tests. Demo audios are publicly available at https://amphionspace.github.io/anyenhance. An open-source implementation is provided at https://github.com/viewfinder-annn/anyenhance-v1-ccf-aatc.

语音增强生成模型提示引导

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。