arXiv:2410.21233cs.SDeess.AS2024-10中稿 · ISMIR 2024被引 28

通过推理时优化,实现对任意音频效果的风格迁移控制。

ST-ITO: Controlling Audio Effects for Style Transfer with Inference-Time Optimization

  • 在推理阶段搜索音频效果链参数空间,无需可微分限制。
  • 自监督预训练构建音频制作风格度量,支持非可微效果控制。
  • 提出多部分评测基准,验证风格迁移表达力与有效性。

音频制作风格迁移旨在将参考录音的风格特征赋予输入音频。现有方法通常训练神经网络以估计一组音频效果的控制参数,但受限于只能控制固定且可微分的效果链,且需特殊训练技术。本文提出ST-ITO(推理时优化的风格迁移),在推理阶段直接搜索音频效果链的参数空间,从而可控制任意效果链,包括未见过的和不可微分的效果。该方法基于一个通过简单可扩展的自监督预训练策略学习的音频制作风格度量,并结合无梯度优化器。由于现有评估手段有限,我们引入一个多部分基准来评估音频风格度量与风格迁移系统。实验表明,所提音频表示更准确捕捉音频制作相关属性,并能通过控制任意效果链实现丰富多样的风格迁移。

原文摘要 · Abstract (English)

Audio production style transfer is the task of processing an input to impart stylistic elements from a reference recording. Existing approaches often train a neural network to estimate control parameters for a set of audio effects. However, these approaches are limited in that they can only control a fixed set of effects, where the effects must be differentiable or otherwise employ specialized training techniques. In this work, we introduce ST-ITO, Style Transfer with Inference-Time Optimization, an approach that instead searches the parameter space of an audio effect chain at inference. This method enables control of arbitrary audio effect chains, including unseen and non-differentiable effects. Our approach employs a learned metric of audio production style, which we train through a simple and scalable self-supervised pretraining strategy, along with a gradient-free optimizer. Due to the limited existing evaluation methods for audio production style transfer, we introduce a multi-part benchmark to evaluate audio production style metrics and style transfer systems. This evaluation demonstrates that our audio representation better captures attributes related to audio production and enables expressive style transfer via control of arbitrary audio effects.

风格迁移音频处理推理优化自监督学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。