arXiv:2512.15313cs.SDcs.LG2025-12中稿 · the Journal of the…

无需控制信号,用GAN建模动态音频效果

Time-Varying Audio Effect Modeling by End-to-End Adversarial Training

  • 用对抗训练+状态预测网络,仅凭音轨建模时变音频效果
  • 在复古移相器上实现高精度时变响应模拟
  • 适合无法获取控制信号的黑箱音频系统建模

深度学习已成为音频效果建模的标准方法,但对时变系统而言,纯黑箱建模仍具挑战。与静态效果不同,含内部调制的设备通常需记录或提取控制信号以保证标准损失函数的时间对齐。本文提出一种生成对抗网络(GAN)框架,仅通过输入输出音频记录即可建模此类效果,无需调制信号提取。采用卷积-循环架构,分两阶段训练:先通过对抗训练使模型学习调制行为分布,不依赖严格相位约束;再通过监督微调,由状态预测网络(SPN)估计初始内部状态,实现模型与目标同步。此外,设计基于啁啾信号的新度量指标,量化调制精度。实验表明,该方法在复古硬件移相器上可有效捕捉时变动态,在完全黑箱条件下表现优异。

原文摘要 · Abstract (English)

Deep learning has become a standard approach for the modeling of audio effects, yet strictly black-box modeling remains problematic for time-varying systems. Unlike time-invariant effects, training models on devices with internal modulation typically requires the recording or extraction of control signals to ensure the time-alignment required by standard loss functions. This paper introduces a Generative Adversarial Network (GAN) framework to model such effects using only input-output audio recordings, without requiring a modulation signal extraction. We propose a convolutional-recurrent architecture trained via a two-stage strategy: an initial adversarial phase allows the model to learn the distribution of the modulation behavior without strict phase constraints, followed by a supervised fine-tuning phase where a State Prediction Network (SPN) estimates the initial internal states required to synchronize the model with the target. Additionally, a new metric based on chirp-train signals is developed to quantify modulation accuracy. Experiments modeling a vintage hardware phaser demonstrate the method's ability to capture time-varying dynamics in a fully black-box context.

音频建模GAN时变系统黑箱建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。