arXiv:2511.14138eess.AScs.SD2025-11被引 4

不用梯度,用搜索找到最符合文字描述的音频效果组合。

FxSearcher: gradient-free text-driven audio transformation

  • 用贝叶斯优化+文本评分器自动找最佳音频处理参数
  • 在多个指标上得分媲美人类偏好,最高达9.2分(满分10)
  • 适合需要精准控制音效风格的音乐创作和影视后期

从文本提示实现多样且高质量的音频变换仍具挑战性,因现有方法严重依赖有限的可微分音频效果。本文提出FxSearcher,一种新型无梯度框架,通过贝叶斯优化与基于CLAP的评分函数,高效搜索满足文本提示的最优音频效果配置。引入引导提示以避免不良失真并提升人听偏好。为客观评估,我们构建了基于AI的评价体系。结果表明,该方法在各项指标上的最高得分与人类偏好高度一致,最高达9.2分(满分10)。演示视频见https://hojoonki.github.io/FxSearcher/

原文摘要 · Abstract (English)

Achieving diverse and high-quality audio transformations from text prompts remains challenging, as existing methods are fundamentally constrained by their reliance on a limited set of differentiable audio effects. This paper proposes FxSearcher, a novel gradient-free framework that discovers the optimal configuration of audio effects (FX) to transform a source signal according to a text prompt. Our method employs Bayesian Optimization and CLAP-based score function to perform this search efficiently. Furthermore, a guiding prompt is introduced to prevent undesirable artifacts and enhance human preference. To objectively evaluate our method, we propose an AI-based evaluation framework. The results demonstrate that the highest scores achieved by our method on these metrics align closely with human preferences. Demos are available at https://hojoonki.github.io/FxSearcher/

音频生成文本控制无梯度优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。