提出实时音乐增强的基准,揭示降噪需考虑音源特性与立体声信息。
Low-Latency Neural Models for Real-Time Music Enhancement

- 设计低延迟因果网络,适配音乐信号的复杂特性。
- 实测所有模型均超实时运行,但效果依赖数据集与失真类型。
- 强调需感知失真类型、保持音色原貌,评价应多维度综合。
音乐录音和直播常受噪声、混响、频谱失衡或伪影影响,降低听感质量。与成熟的语音增强不同,音乐增强因信号包含重叠音源、宽频带、强动态及有意制作效果而更具挑战。本文在严格因果与低延迟约束下研究实时音乐增强,聚焦从声学与制作层面退化中恢复原始混音。通过适配紧凑因果网络,并对比语音基线、外部音乐去噪模型、离线修复参考以及音乐专用的MusicFilterNet-MS变体。在测试硬件上,所有因果模型均运行于实时速度以上,但性能显著依赖数据集、退化类型与评估指标;多个客观标准下,盲目增强反而可能恶化输入。核心贡献在于建立基准与分析框架:实时音乐增强可行,但鲁棒提升需依赖退化感知建模、立体声感知处理、音色保留修正,以及超越单一指标的综合评估。
原文摘要 · Abstract (English)
Music recordings and live streams are often affected by noise, reverberation, spectral imbalances, or artifacts that degrade listening quality. While speech enhancement has matured into a well-defined research area, music enhancement is less established because musical signals combine overlapping sources, wide bandwidths, strong dynamics, and intentional production effects. We study real-time music enhancement under strict causal and low-latency constraints. We formulate the task around recovery of the intended produced mix from acoustic and production-oriented degradations, adapt compact causal networks to music, and compare speech-derived real-time baselines, an external music-denoising model, an offline restoration reference, and a music-specific MusicFilterNet-MS variant. On the tested hardware, all causal models run faster than real time, but improvements depend strongly on the dataset, degradation type, and metric family; under several objective criteria, indiscriminate enhancement can worsen the degraded input. The main contribution is therefore a benchmark and an analysis rather than a universal best model: real-time music enhancement is feasible, but robust improvement requires degradation-aware modeling, stereo-aware processing, identity-preserving correction, and evaluation beyond a single objective score.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。