用深度神经网络预测干信号和音效参数,再通过重构相似性优化结果。
Audio Effect Estimation with DNN-Based Prediction and Search Algorithm
- 先用DNN预测干信号与音效配置,再基于湿信号重构进行搜索优化。
- 新方法在多个指标上优于纯预测方法,尤其在音效顺序与参数估计上表现更好。
- 适合音效设计、音频修复领域研究者参考,对音效链逆向分析有实用价值。
音效在声音设计中至关重要。本文研究从湿信号中估计所应用音效的配置问题。现有方法分为预测型(数据驱动训练模型)和搜索型(基于湿信号重构)。本研究提出一种融合方法:首先用深度神经网络预测干信号与音效配置,然后基于这些预测进行湿信号重构搜索。通过在预测阶段估计干信号,可利用重构相似性作为目标函数来补充或改进预测。实验表明,该方法在多种评估指标上优于仅依赖预测的方法。此外,先预测音效类型组合,再搜索顺序与参数的分步策略表现最优。
原文摘要 · Abstract (English)
Audio effects play an essential role in sound design. This research addresses the task of audio effect estimation, which aims to estimate the configuration of applied effects from a wet signal. Existing approaches to this problem can be categorized into predictive approaches, which use models pre-trained in a data-driven manner, and search-based approaches, which are based on wet signal reconstruction. In this study, we propose a novel approach that integrates these approaches: first, DNNs predict the dry signal and effect configuration, and then a search is performed based on wet signal reconstruction using these predictions. By estimating the dry signal in the prediction stage, it becomes possible to complement or improve the predictions using reconstruction similarity as an objective function. The experimental evaluation showed that methods based on the proposed approach outperformed the method solely based on the predictive approach. Furthermore, the findings suggest that the task division of predicting the effect type combination followed by the search-based estimation of order and parameters was the most effective across various metrics.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。