对比多种声学相似性度量在不同合成器上的表现,发现无通用最优方案。
Evaluating Sound Similarity Metrics for Differentiable, Iterative Sound-Matching
- 设计可微分迭代音色匹配框架,结合机器学习与人工调音流程。
- 16种合成器-损失组合共300次实验,结果依赖具体合成方法。
- 建议针对特定合成技术定制相似性度量,而非追求通用解。
手工音色设计本质上是迭代过程:艺术家将合成输出与内心目标比较,调整参数并重复直至满意。迭代音色匹配通过损失函数(或相似性度量)指导合成器持续编程,自动化此流程。以往对损失函数的比较多局限于有限的合成方法和少量损失类型,且常缺乏盲听测试。这使得是否存有通用最优损失函数仍不明朗,或选择损失仍取决于合成方法与设计师偏好。本文提出可微分迭代音色匹配,作为现有研究的自然延伸,融合机器学习进展与人工设计流程。为分析损失函数在不同合成器上的表现差异,我们实现四种新型及经典可微分损失函数,并与可微分减法、加法及AM合成器配对。对每组16种合成器-损失组合,执行300次随机音色匹配试验,以参数差异、频谱距离指标和人工评分衡量性能。三类评估指标间呈现中等一致性。后验分析表明,损失函数表现高度依赖于合成器。该结果强调拓展音色匹配实验范围的重要性,并支持针对特定合成技术开发定制化相似性度量,而非追求通用解。
原文摘要 · Abstract (English)
Manual sound design with a synthesizer is inherently iterative: an artist compares the synthesized output to a mental target, adjusts parameters, and repeats until satisfied. Iterative sound-matching automates this workflow by continually programming a synthesizer under the guidance of a loss function (or similarity measure) toward a target sound. Prior comparisons of loss functions have typically favored one metric over another, but only within narrow settings: limited synthesis methods, few loss types, often without blind listening tests. This leaves open the question of whether a universally optimal loss exists, or the choice of loss remains a creative decision conditioned on the synthesis method and the sound designer's preference. We propose differentiable iterative sound-matching as the natural extension of the available literature, since it combines the manual approach to sound design with modern advances in machine learning. To analyze the variability of loss function performance across synthesizers, we implemented a mix of four novel and established differentiable loss functions, and paired them with differentiable subtractive, additive, and AM synthesizers. For each of the sixteen synthesizer--loss combinations, we ran 300 randomized sound-matching trials. Performance was measured using parameter differences, spectrogram-distance metrics, and manually assigned listening scores. We observed a moderate level of consistency among the three performance measures. Our post-hoc analysis shows that the loss function performance is highly dependent on the synthesizer. These findings underscore the value of expanding the scope of sound-matching experiments and developing new similarity metrics tailored to specific synthesis techniques rather than pursuing one-size-fits-all solutions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。