arXiv:2509.24286eess.AScs.SD2025-09被引 1

分离音色、包络与内容,实现可调控的电子合成器音效迁移。

SynthCloner: Synthesizer-style Audio Transfer via Factorized Codec with ADSR Envelope Control

  • 将音频分解为包络、音色和内容三要素,独立控制合成器音效
  • 在250种音色、120种包络上验证,主观评测得分超越基线
  • 开源数据集含100段MIDI,适合音效设计与音乐生成研究者

电子合成器音效由参数设定决定,产生复杂的音色特征与包络形态,使合成器风格音频迁移尤为困难。现有方法多依赖频谱目标或隐式风格匹配,难以精细控制包络。此外,公开合成器数据集往往缺乏多样化的音色与包络覆盖。为此,我们提出SynthCloner,一种因子化编解码模型,将音频解耦为三类属性:ADSR包络、音色与内容,实现对各属性的独立调控。同时,我们构建了新型合成器数据集SynthCAT,采用任务特定渲染流程,覆盖250种音色、120种ADSR包络及100段MIDI序列。实验表明,SynthCloner在客观与主观评价上均优于基线方法,并支持灵活属性控制。代码、模型检查点与音频示例已开源。

原文摘要 · Abstract (English)

Electronic synthesizer sounds are controlled by parameter settings that yield complex timbral characteristics and ADSR envelopes, making synthesizer-style audio transfer particularly challenging. Recent approaches to timbre transfer often rely on spectral objectives or implicit style matching, offering limited control over envelope shaping. Moreover, public synthesizer datasets rarely provide diverse coverage of timbres and ADSR envelopes. To address these gaps, we present SynthCloner, a factorized codec model that disentangles audio into three attributes: ADSR envelope, timbre, and content. This separation enables expressive audio transfer with independent control over these attributes. Additionally, we introduce SynthCAT, a new synthesizer dataset with a task-specific rendering pipeline covering 250 timbres, 120 ADSR envelopes, and 100 MIDI sequences. Experiments show that SynthCloner outperforms baselines on both objective and subjective metrics, while enabling independent attribute control. The code, model checkpoint, and audio examples are available at https://buffett0323.github.io/synthcloner/.

音频迁移音色控制合成器因子化编码

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。