arXiv:2504.04751eess.AScs.AI2025-04中稿 · the 28th Internati…被引 5

用扩散模型无监督估计音频失真效果,更稳定且对数据量要求低。

Unsupervised Estimation of Nonlinear Audio Effects: Comparing Diffusion-Based and Adversarial approaches

  • 采用扩散生成模型实现无监督音频失真盲识别,支持黑箱与灰箱建模。
  • 在吉他失真实验中,扩散方法更稳定,数据少时表现更优。
  • 适合音乐科技中未知音频效果的鲁棒估计,尤其适合数据不足场景。

在无法获取配对输入输出信号的情况下,准确估计非线性音频效果仍是一项挑战。本文研究了无监督概率方法解决该问题。提出一种新颖的基于扩散生成模型的盲系统识别方法,可利用黑箱或灰箱模型估计未知非线性效应。对比此前提出的对抗式方法,在不同参数设置和可用录音长度条件下评估两者的性能。通过对吉他失真效果的实验表明,扩散方法结果更稳定,对数据可用性不敏感;而对抗方法在估计明显失真时表现更优。研究为音频效果的鲁棒无监督盲估计提供了新思路,展示了扩散模型在音乐技术中系统识别的潜力。

原文摘要 · Abstract (English)

Accurately estimating nonlinear audio effects without access to paired input-output signals remains a challenging problem. This work studies unsupervised probabilistic approaches for solving this task. We introduce a method, novel for this application, based on diffusion generative models for blind system identification, enabling the estimation of unknown nonlinear effects using black- and gray-box models. This study compares this method with a previously proposed adversarial approach, analyzing the performance of both methods under different parameterizations of the effect operator and varying lengths of available effected recordings. Through experiments on guitar distortion effects, we show that the diffusion-based approach provides more stable results and is less sensitive to data availability, while the adversarial approach is superior at estimating more pronounced distortion effects. Our findings contribute to the robust unsupervised blind estimation of audio effects, demonstrating the potential of diffusion models for system identification in music technology.

音频处理扩散模型无监督学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。