用可微信号处理技术,自动发现音频中复杂的调制模式。
Modulation Discovery with Differentiable Digital Signal Processing
- 基于可微数字信号处理,从音频中反推调制信号的形状与结构。
- 在合成音与真实音频上实现高精度匹配,同时保持调制过程可解释。
- 适合音乐制作人与音频工程师,用于分析和复现复杂音色设计。
调制是声音设计与音乐制作的核心,能生成复杂多变的音频效果。现代合成器提供包络、低频振荡器(LFO)及参数自动化工具,但难以逆向解析声音背后的调制信号。现有音色匹配或参数估计方法常为不可解释的黑箱,或仅预测高维帧级参数,忽略调制曲线的形态、结构与路由关系。本文提出一种神经音色匹配方法,结合调制提取、约束控制信号参数化与可微数字信号处理(DDSP),以发现音频中实际存在的调制机制。我们在高度调制的合成音与真实音频样本上验证了该方法的有效性,展示了其对不同DDSP合成架构的适用性,并研究了可解释性与音色匹配精度之间的权衡。代码与音频样本已公开,并提供可运行的VST插件版本。
原文摘要 · Abstract (English)
Modulations are a critical part of sound design and music production, enabling the creation of complex and evolving audio. Modern synthesizers provide envelopes, low frequency oscillators (LFOs), and more parameter automation tools that allow users to modulate the output with ease. However, determining the modulation signals used to create a sound is difficult, and existing sound-matching / parameter estimation systems are often uninterpretable black boxes or predict high-dimensional framewise parameter values without considering the shape, structure, and routing of the underlying modulation curves. We propose a neural sound-matching approach that leverages modulation extraction, constrained control signal parameterizations, and differentiable digital signal processing (DDSP) to discover the modulations present in a sound. We demonstrate the effectiveness of our approach on highly modulated synthetic and real audio samples, its applicability to different DDSP synth architectures, and investigate the trade-off it incurs between interpretability and sound-matching accuracy. We make our code and audio samples available and provide the trained DDSP synths in a VST plugin.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。