用语义嵌入实现音频自动均衡,效果接近理想情况。
Automatic Audio Equalization with Semantic Embeddings

- 用预训练模型提取语义特征,轻量头部训练提升效率。
- 客观评估有效,主观测试媲美理想参考模型。
- 适合需要实时音频增强的场景,尤其抗噪声与混响。
本文提出一种数据驱动的自动盲音频均衡方法,通过预测对数梅尔频谱特征并推导逆滤波器实现。该方法采用深度神经网络,以预训练模型提供语义嵌入作为主干,仅训练轻量级头部,旨在提升训练效率与泛化能力。模型在音乐和语音数据上训练,对噪声和混响具有鲁棒性。客观评估证实其有效性,主观测试显示性能接近使用真实对数梅尔频谱特征的最优参考模型(oracle),表明模型能准确估计所需特性,剩余局限主要归因于滤波阶段。整体结果展示了该方法在真实音频增强应用中的潜力。
原文摘要 · Abstract (English)
This paper presents a data-driven approach to automatic blind equalization of audio by predicting log-mel spectral features and deriving an inverse filter. The method uses a deep neural network, where a pre-trained model provides semantic embeddings as a backbone, and only a lightweight head is trained. This design is intended to enhance training efficiency and generalization. Trained on both music and speech, the model is robust to noise and reverberation. Objective evaluations confirm its effectiveness, and subjective tests show performance comparable to that of an oracle that uses true log-mel spectral features, indicating that the model accurately estimates the desired characteristics, with remaining limitations attributed to the filtering stage. Overall, the results highlight the potential of the method for real-world audio enhancement applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。