arXiv:2409.18847eess.AScs.SD2024-09中稿 · ICASSP 2025被引 19

用自然语言控制音频效果,无需重训练模型。

Text2FX: Harnessing CLAP Embeddings for Text-Guided Audio Effects

  • 通过CLAP嵌入空间优化,将文本转为音效参数。
  • 支持“在你面前”等开放词汇指令,实现音效精准调节。
  • 适用于各类可微分音效,适合音视频创作与交互设计。

本文提出Text2FX,利用CLAP嵌入和可微分数字信号处理,通过开放式自然语言提示(如“让声音显得有冲击力且大胆”)控制均衡器和混响等音频效果。该方法无需重新训练模型,仅在现有嵌入空间中进行单实例优化,实现灵活、可扩展的开放词汇音效变换。实验表明,CLAP编码了控制音效的有效信息,并提出两种基于CLAP的优化策略以将文本映射为音效参数。该方法不仅限于CLAP,也适用于任何共享的文本-音频嵌入空间;同样不限于均衡与混响,可拓展至任意可微分音效。我们通过包含多样文本提示和源音频的听觉测试评估了方法在人类感知上的质量与对齐度。演示与代码已公开于anniejchu.github.io/text2fx。

原文摘要 · Abstract (English)

This work introduces Text2FX, a method that leverages CLAP embeddings and differentiable digital signal processing to control audio effects, such as equalization and reverberation, using open-vocabulary natural language prompts (e.g., "make this sound in-your-face and bold"). Text2FX operates without retraining any models, relying instead on single-instance optimization within the existing embedding space, thus enabling a flexible, scalable approach to open-vocabulary sound transformations through interpretable and disentangled FX manipulation. We show that CLAP encodes valuable information for controlling audio effects and propose two optimization approaches using CLAP to map text to audio effect parameters. While we demonstrate with CLAP, this approach is applicable to any shared text-audio embedding space. Similarly, while we demonstrate with equalization and reverberation, any differentiable audio effect may be controlled. We conduct a listener study with diverse text prompts and source audio to evaluate the quality and alignment of these methods with human perception. Demos and code are available at anniejchu.github.io/text2fx.

音频生成文本控制CLAP可微分音效

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。