arXiv:2505.04621cs.SDcs.AI2025-05被引 5

用扩散模型生成音频,能分离声源、模拟音效、调参数。

Score Distillation Sampling for Audio: Source Separation, Synthesis, and Beyond

  • 用文本引导的音频扩散模型,通过蒸馏生成先验知识。
  • 单个预训练模型完成音效模拟、参数校准和声源分离。
  • 适合想用生成模型做音频处理的研究者和开发者。

我们提出Audio-SDS,将得分蒸馏采样(Score Distillation Sampling, SDS)扩展至文本条件音频扩散模型。尽管SDS最初用于文本到3D生成,其核心思想——将强大生成先验蒸馏为独立参数化表示——可推广至音频领域。借助单一预训练模型,Audio-SDS无需专用数据集即可实现多种任务:指导物理启发的撞击声模拟、校准FM合成参数、执行提示指定的声源分离。结果表明,基于蒸馏的方法在多模态中具有高度通用性,为未来利用生成先验解决音频任务奠定坚实基础。

原文摘要 · Abstract (English)

We introduce Audio-SDS, a generalization of Score Distillation Sampling (SDS) to text-conditioned audio diffusion models. While SDS was initially designed for text-to-3D generation using image diffusion, its core idea of distilling a powerful generative prior into a separate parametric representation extends to the audio domain. Leveraging a single pretrained model, Audio-SDS enables a broad range of tasks without requiring specialized datasets. In particular, we demonstrate how Audio-SDS can guide physically informed impact sound simulations, calibrate FM-synthesis parameters, and perform prompt-specified source separation. Our findings illustrate the versatility of distillation-based methods across modalities and establish a robust foundation for future work using generative priors in audio tasks.

音频生成扩散模型声源分离

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。