arXiv:2409.08723eess.AS2024-09被引 10

FLAMO让音频处理模块可微分,便于在神经网络中优化音频效果。

FLAMO: An Open-Source Library for Frequency-Domain Differentiable Audio Processing

  • 基于频域采样法构建可微分音频模块,支持端到端训练
  • 实现混响器与主动声学系统的响应色彩优化,提升音质表现
  • 开源库含预置模块与训练工具,适合音频算法研发者使用

我们提出FLAMO,一个面向音频模块优化的频域采样库,用于实现和优化可微分的线性时不变音频系统。该库开源,基于频域采样滤波器设计方法,可创建独立或嵌入神经网络计算图中的可微分模块,简化可微分音频系统开发。库内包含预定义滤波模块及构建、训练、日志记录辅助类,通过直观接口访问。通过两个案例展示实际应用:人工混响器优化与主动声学系统改进,有效减少响应色偏,提升音质。

原文摘要 · Abstract (English)

We present FLAMO, a Frequency-sampling Library for Audio-Module Optimization designed to implement and optimize differentiable linear time-invariant audio systems. The library is open-source and built on the frequency-sampling filter design method, allowing for the creation of differentiable modules that can be used stand-alone or within the computation graph of neural networks, simplifying the development of differentiable audio systems. It includes predefined filtering modules and auxiliary classes for constructing, training, and logging the optimized systems, all accessible through an intuitive interface. Practical application of these modules is demonstrated through two case studies: the optimization of an artificial reverberator and an active acoustics system for improved response coloration.

音频处理可微分开源库频域

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。