用神经网络+模型方法,精准还原被压缩音频的原始动态范围。
Neural-Enhanced Dynamic Range Compression Inversion: A Hybrid Approach for Restoring Audio Dynamics
- 结合神经网络与物理模型,同步估计压缩参数并还原音频
- 在多个音乐和语音数据集上优于现有顶尖方法
- 适合需要高质量音频修复的音乐制作与广播场景
动态范围压缩(DRC)是音乐制作、广播和语音处理中广泛使用的音频效果。逆向恢复DRC对还原原始音频动态、支持重混音和提升音质具有重要意义。现有方法或忽略关键参数,或依赖精确的参数值,而这些值难以准确估计。为此,我们提出一种混合方法,将基于模型的DRC逆向技术与神经网络相结合,实现鲁棒的参数估计与音频重建。采用定制化的神经网络架构(分类与回归),集成至模型驱动的逆向框架中以恢复原始信号。在多种音乐和语音数据集上的实验验证了该方法的有效性与鲁棒性,性能优于多个先进基准方法。
原文摘要 · Abstract (English)
Dynamic Range Compression (DRC) is a widely used audio effect that adjusts signal dynamics for applications in music production, broadcasting, and speech processing. Inverting DRC is of broad importance for restoring the original dynamics, enabling remixing, and enhancing the overall audio quality. Existing DRC inversion methods either overlook key parameters or rely on precise parameter values, which can be challenging to estimate accurately. To address this limitation, we introduce a hybrid approach that combines model-based DRC inversion with neural networks to achieve robust DRC parameter estimation and audio restoration simultaneously. Our method uses tailored neural network architectures (classification and regression), which are then integrated into a model-based inversion framework to reconstruct the original signal. Experimental evaluations on various music and speech datasets confirm the effectiveness and robustness of our approach, outperforming several state-of-the-art techniques.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。