arXiv:2410.02144cs.SDcs.LG2024-10被引 8

用扩散模型生成听感均匀的音色渐变,让声音过渡更自然。

SoundMorpher: Perceptually-Uniform Sound Morphing with Diffusion Model

  • 基于对数梅尔频谱,显式建模音色变化与感知之间的关系。
  • 通过二分搜索确保每步变化听感一致,实现稳定渐变轨迹。
  • 提出新评估指标,支持客观对比不同音色变换方法。

我们提出 SoundMorpher,一种面向开放世界的音色渐变方法,旨在生成听觉上均匀的渐变路径。传统音色渐变方法通常假设渐变参数与听觉感知呈线性关系,通过线性插值源音与目标音的语义特征来实现平滑过渡。然而,这类方法过度简化了声音感知的复杂性,导致渐变质量受限。相比之下,SoundMorpher 显式建模渐变参数与感知之间的关系,利用对数梅尔频谱特征,通过二分搜索确定使每一步听觉差异恒定的渐变参数,从而优化渐变序列。为解决音色渐变缺乏定量评估框架的问题,我们提出了基于三个客观标准的评估指标,可全面评估生成结果并实现方法间的直接比较,推动该领域发展。大量实验表明,SoundMorpher 在真实场景中表现优异,适用于音乐创作、影视后期及交互音频等应用。演示与代码已公开于~\url{https://xinleiniu.github.io/SoundMorpher-demo/}。

原文摘要 · Abstract (English)

We present SoundMorpher, an open-world sound morphing method designed to generate perceptually uniform morphing trajectories. Traditional sound morphing techniques typically assume a linear relationship between the morphing factor and sound perception, achieving smooth transitions by linearly interpolating the semantic features of source and target sounds while gradually adjusting the morphing factor. However, these methods oversimplify the complexities of sound perception, resulting in limitations in morphing quality. In contrast, SoundMorpher explores an explicit relationship between the morphing factor and the perception of morphed sounds, leveraging log Mel-spectrogram features. This approach further refines the morphing sequence by ensuring a constant target perceptual difference for each transition and determining the corresponding morphing factors using binary search. To address the lack of a formal quantitative evaluation framework for sound morphing, we propose a set of metrics based on three established objective criteria. These metrics enable comprehensive assessment of morphed results and facilitate direct comparisons between methods, fostering advancements in sound morphing research. Extensive experiments demonstrate the effectiveness and versatility of SoundMorpher in real-world scenarios, showcasing its potential in applications such as creative music composition, film post-production, and interactive audio technologies. Our demonstration and codes are available at~\url{https://xinleiniu.github.io/SoundMorpher-demo/}.

音色变换扩散模型听觉感知

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。