通过等变流匹配解决音色合成器参数反演的对称性难题
Audio synthesizer inversion in symmetric parameter spaces with approximately equivariant flow matching
- 将参数反演视为概率分布建模,利用条件生成模型处理对称性
- 在Surge XT上实现更优的音频重建效果,峰值信噪比提升显著
- 适用于需要精准音色参数还原的音乐制作与音频分析场景
许多音频合成器在不同参数配置下可生成相同信号,导致从声音反推参数本质上是病态问题。我们发现这主要源于合成器内在对称性,特别是置换不变性。首先,在合成任务中表明,即使使用置换不变损失函数或对称性破坏启发式方法,基于点估计的回归性能仍会下降。随后,将等价解视为概率分布中的模式,采用条件生成模型显著提升性能。进一步考虑到隐式参数分布的不变性,使用置换等变连续归一化流进一步优化。为应对真实合成器中的复杂对称性,提出一种自适应发现相关对称性的松弛等变策略。在开源全功能合成器Surge XT上的实验显示,该方法在音频重建指标上优于回归与生成基线。
原文摘要 · Abstract (English)
Many audio synthesizers can produce the same signal given different parameter configurations, meaning the inversion from sound to parameters is an inherently ill-posed problem. We show that this is largely due to intrinsic symmetries of the synthesizer, and focus in particular on permutation invariance. First, we demonstrate on a synthetic task that regressing point estimates under permutation symmetry degrades performance, even when using a permutation-invariant loss function or symmetry-breaking heuristics. Then, viewing equivalent solutions as modes of a probability distribution, we show that a conditional generative model substantially improves performance. Further, acknowledging the invariance of the implicit parameter distribution, we find that performance is further improved by using a permutation equivariant continuous normalizing flow. To accommodate intricate symmetries in real synthesizers, we also propose a relaxed equivariance strategy that adaptively discovers relevant symmetries from data. Applying our method to Surge XT, a full-featured open source synthesizer used in real world audio production, we find our method outperforms regression and generative baselines across audio reconstruction metrics.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。