从音频自动还原合成器参数,让声音还原更真实。
INSTRUMENTAL: Automatic Synthesizer Parameter Recovery from Audio via Evolutionary Optimization
- 用可微分的28参数合成器配合进化优化,逆向推导音频中的参数。
- 在真实录音上匹配损失低至2.09,效果优于传统梯度下降。
- 发现频段均衡增强能提升收敛,随机初始化不如谱分析启动。
现有音符提取工具仅捕捉音符信息,丢失定义乐器音色的关键特征。本文提出Instrumental系统,通过将可微分的28参数减法合成器与无导数进化优化器CMA-ES结合,实现从音频中恢复连续合成器参数。采用融合梅尔刻度STFT、频谱中心和MFCC差异的复合感知损失函数,在真实录制音频上达到2.09的匹配损失。我们系统评估了八种提升收敛性的假设,发现仅有参数化均衡增强带来显著改善。结果表明:在非凸优化空间中,CMA-ES优于梯度下降;参数数量增加不单调提升匹配性能;基于谱分析的初始化比随机初始化更快收敛。
原文摘要 · Abstract (English)
Existing audio-to-MIDI tools extract notes but discard the timbral characteristics that define an instrument's identity. We present Instrumental, a system that recovers continuous synthesizer parameters from audio by coupling a differentiable 28-parameter subtractive synthesizer with CMA-ES, a derivative-free evolutionary optimizer. We optimize a composite perceptual loss combining mel-scaled STFT, spectral centroid, and MFCC divergence, achieving a matching loss of 2.09 on real recorded audio. We systematically evaluate eight hypotheses for improving convergence and find that only parametric EQ boosting yields meaningful improvement. Our results show that CMA-ES outperforms gradient descent on this non-convex landscape, that more parameters do not monotonically improve matching, and that spectral analysis initialization accelerates convergence over random starts.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。