arXiv:2502.17527cs.SDcs.AI2025-02被引 1

用深度学习优化音乐频谱,让耳机音乐更好掩盖环境噪音

Perceptual Noise-Masking with Music through Deep Spectral Envelope Shaping

  • 基于听觉掩蔽原理,用神经网络重塑音乐频谱包络
  • 在模拟场景下显著提升噪声掩蔽效果,优于现有方法
  • 适合耳机降噪、音频增强等需要听觉感知优化的场景

人们常在嘈杂环境中听音乐,希望隔绝环境声。由于同时掩蔽效应,音乐信号可遮蔽部分噪声频率成分。本文提出一种基于心理声学掩蔽模型的神经网络,通过预测滤波器频率响应来重塑音乐的频谱包络,增强其掩蔽能力。模型采用感知损失函数,在有效掩蔽噪声的同时,保持原始音乐混音和用户设定的音量水平。我们在模拟用户佩戴耳机在嘈杂环境中听音乐的场景下评估该方法,基于客观指标的结果表明,该系统优于当前最先进方法。

原文摘要 · Abstract (English)

People often listen to music in noisy environments, seeking to isolate themselves from ambient sounds. Indeed, a music signal can mask some of the noise's frequency components due to the effect of simultaneous masking. In this article, we propose a neural network based on a psychoacoustic masking model, designed to enhance the music's ability to mask ambient noise by reshaping its spectral envelope with predicted filter frequency responses. The model is trained with a perceptual loss function that balances two constraints: effectively masking the noise while preserving the original music mix and the user's chosen listening level. We evaluate our approach on simulated data replicating a user's experience of listening to music with headphones in a noisy environment. The results, based on defined objective metrics, demonstrate that our system improves the state of the art.

音频增强听觉掩蔽神经网络

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。