用语音样例类比实现声音纹理的精准操控。
Audio Texture Manipulation by Exemplar-Based Analogy
- 通过配对语音样例学习声音变换,替代文本指令。
- 在多个编辑任务上超越文本基基线模型,泛化能力强。
- 适合需要灵活声音编辑的音频生成与后期制作场景。
音频纹理操纵旨在修改声音的感知特征以实现特定变换,如添加、移除或替换听觉元素。本文提出一种基于样例类比的音频纹理操纵模型。不同于依赖文本指令的条件输入,该方法使用成对语音片段:一个代表原始声音,另一个展示期望的变换效果。模型学习将相同变换应用于新输入,从而实现声音纹理的操控。我们构建了一个四元组数据集,涵盖多种编辑任务,并以自监督方式训练潜空间扩散模型。定量评估和感知研究表明,该模型优于文本条件基线,在真实世界、分布外及非语音场景下均具备良好泛化能力。
原文摘要 · Abstract (English)
Audio texture manipulation involves modifying the perceptual characteristics of a sound to achieve specific transformations, such as adding, removing, or replacing auditory elements. In this paper, we propose an exemplar-based analogy model for audio texture manipulation. Instead of conditioning on text-based instructions, our method uses paired speech examples, where one clip represents the original sound and another illustrates the desired transformation. The model learns to apply the same transformation to new input, allowing for the manipulation of sound textures. We construct a quadruplet dataset representing various editing tasks, and train a latent diffusion model in a self-supervised manner. We show through quantitative evaluations and perceptual studies that our model outperforms text-conditioned baselines and generalizes to real-world, out-of-distribution, and non-speech scenarios. Project page: https://berkeley-speech-group.github.io/audio-texture-analogy/
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。