用相似度向量实现音效的精细可控合成,让声音变化更自然。
Simi-SFX: A similarity-based conditioning method for controllable sound effect synthesis
- 基于可微信号处理,用归一化向量控制音色特征。
- 在两个新数据集上验证,相似度分数能精准调控音色变化。
- 适合音效设计、音乐生成等需要细腻控制的场景。
可控音效生成是一项挑战性任务,传统方法依赖复杂的物理模型,需深入了解信号处理参数。随着生成模型的发展,文本成为常见的控制接口,但语言符号的离散性和定性特征难以捕捉声音间的细微音色差异。本文提出一种基于相似性的音效合成条件方法,结合可微数字信号处理(DDSP),利用潜在空间学习和控制音频音色,并引入一个归一化范围为[0,1]的引导向量来编码类别声学信息。通过预训练音频表示模型,该方法实现了丰富且细粒度的音色控制。为评估方法,我们构建了两个音效数据集——Footstep-set 和 Impact-set,用于衡量可控性与音质。回归分析表明,所提出的相似度分数能有效控制音色变化,并支持音色插值等创意应用。本研究提供了一个强大且通用的音效合成框架,弥合了传统信号处理与现代机器学习技术之间的差距。
原文摘要 · Abstract (English)
Generating sound effects with controllable variations is a challenging task, traditionally addressed using sophisticated physical models that require in-depth knowledge of signal processing parameters and algorithms. In the era of generative and large language models, text has emerged as a common, human-interpretable interface for controlling sound synthesis. However, the discrete and qualitative nature of language tokens makes it difficult to capture subtle timbral variations across different sounds. In this research, we propose a novel similarity-based conditioning method for sound synthesis, leveraging differentiable digital signal processing (DDSP). This approach combines the use of latent space for learning and controlling audio timbre with an intuitive guiding vector, normalized within the range [0,1], to encode categorical acoustic information. By utilizing pre-trained audio representation models, our method achieves expressive and fine-grained timbre control. To benchmark our approach, we introduce two sound effect datasets--Footstep-set and Impact-set--designed to evaluate both controllability and sound quality. Regression analysis demonstrates that the proposed similarity score effectively controls timbre variations and enables creative applications such as timbre interpolation between discrete classes. Our work provides a robust and versatile framework for sound effect synthesis, bridging the gap between traditional signal processing and modern machine learning techniques.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。