arXiv:2601.20426cs.SD2026-01中稿 · to ICASSP 2026被引 1

用噪声混合训练音频变形模型,实现高质量音色融合。

Mix2Morph: Learning Sound Morphing from Noisy Mixes

  • 在高扩散步骤的噪声混合上微调,无需专门的变形数据集。
  • 生成的音频融合自然,感知上连贯,能保留双源特征。
  • 适合音效设计、音乐创作等需要可控声学融合的场景。

我们提出 Mix2Morph,一个基于文本到音频扩散模型的微调方法,可在无专用变形数据集的情况下实现声音形态变换。通过在较高扩散时间步的噪声代理混合数据上进行微调,该模型生成稳定且感知一致的声音融合效果,能够自然整合两个声音源的特性。研究聚焦于声音注入这一实际且感知有意义的子类:一个声音作为主导主源,决定整体时序与结构行为,另一个声音被持续注入以丰富其音色与质感。客观评估和听觉测试表明,Mix2Morph优于现有基线,在多种声音类别中均产生高质量的声音注入效果,为更可控、概念驱动的声音设计工具迈进一步。音频示例见 https://anniejchu.github.io/mix2morph。

原文摘要 · Abstract (English)

We introduce Mix2Morph, a text-to-audio diffusion model fine-tuned to perform sound morphing without a dedicated dataset of morphs. By finetuning on noisy surrogate mixes at higher diffusion timesteps, Mix2Morph yields stable, perceptually coherent morphs that convincingly integrate qualities of both sources. We specifically target sound infusions, a practically and perceptually motivated subclass of morphing in which one sound acts as the dominant primary source, providing overall temporal and structural behavior, while a secondary sound is infused throughout, enriching its timbral and textural qualities. Objective evaluations and listening tests show that Mix2Morph outperforms prior baselines and produces high-quality sound infusions across diverse categories, representing a step toward more controllable and concept-driven tools for sound design. Sound examples are available at https://anniejchu.github.io/mix2morph .

音频生成声音融合扩散模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。