用专业音频软件生成复杂音效链数据,让AI更懂真实音乐制作
WildFX: A DAW-Powered Pipeline for In-the-Wild Audio FX Graph Modeling
- 基于DAW的容器化管道,支持多种音效插件和复杂信号流
- 在盲测中准确还原音效链结构与参数,性能优于简化神经控制器
- 适合研究音频处理建模或想接轨工业级音效流程的AI开发者
尽管端到端AI音乐生成进展迅速,但对专业数字信号处理(DSP)工作流的建模仍具挑战。现有神经黑箱方法难以复现专业工作流中的精细信号路径与参数交互。当前可微分插件方案常偏离真实工具,性能逊于简化神经控制器。本文提出WildFX,一个基于Docker的容器化管道,利用专业数字音频工作站(DAW)后端生成多轨混音数据集,支持跨平台商业插件及任意格式(VST/VST3/LV2/CLAP)插件,实现侧链、分频等复杂结构,并高效并行处理。通过极简元数据接口简化配置。实验表明,该管道在盲测中成功估计混音图结构、插件与增益参数,有效弥合了AI研究与实际DSP需求之间的差距。代码已开源:https://github.com/IsaacYQH/WildFX。
原文摘要 · Abstract (English)
Despite rapid progress in end-to-end AI music generation, AI-driven modeling of professional Digital Signal Processing (DSP) workflows remains challenging. In particular, while there is growing interest in neural black-box modeling of audio effect graphs (e.g. reverb, compression, equalization), AI-based approaches struggle to replicate the nuanced signal flow and parameter interactions used in professional workflows. Existing differentiable plugin approaches often diverge from real-world tools, exhibiting inferior performance relative to simplified neural controllers under equivalent computational constraints. We introduce WildFX, a pipeline containerized with Docker for generating multi-track audio mixing datasets with rich effect graphs, powered by a professional Digital Audio Workstation (DAW) backend. WildFX supports seamless integration of cross-platform commercial plugins or any plugins in the wild, in VST/VST3/LV2/CLAP formats, enabling structural complexity (e.g., sidechains, crossovers) and achieving efficient parallelized processing. A minimalist metadata interface simplifies project/plugin configuration. Experiments demonstrate the pipeline's validity through blind estimation of mixing graphs, plugin/gain parameters, and its ability to bridge AI research with practical DSP demands. The code is available on: https://github.com/IsaacYQH/WildFX.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。