用机器学习从3个点的噪声数据恢复高精度听觉模型
A Machine Learning Approach for Denoising and Upsampling HRTFs
- 先用U-Net去噪,再用AE-GAN从3个点补全完整HRTF
- LSD误差5.41 dB,余弦相似度损失0.0070,效果显著
- 适合快速采集个性化声场数据的研究与应用
虚拟沉浸式音频对真实感的需求持续增长,头相关传递函数(HRTFs)在此中起关键作用。HRTFs反映声音到达双耳的特性,体现个体解剖差异,提升空间感知。研究表明个性化HRTF可提高定位准确率,但测量耗时且需无噪声环境。尽管机器学习能减少所需测量点数以缩短时间,仍依赖受控环境。本文提出一种新方法,通过结合HRTF去噪U-Net与自编码生成对抗网络(AE-GAN),实现从仅三个测量点的稀疏、噪声数据中完成超分辨率重建。该方法在测试中达到5.41 dB的对数谱失真(LSD)误差和0.0070的余弦相似度损失,验证了其在HRTF上采样中的有效性。
原文摘要 · Abstract (English)
The demand for realistic virtual immersive audio continues to grow, with Head-Related Transfer Functions (HRTFs) playing a key role. HRTFs capture how sound reaches our ears, reflecting unique anatomical features and enhancing spatial perception. It has been shown that personalized HRTFs improve localization accuracy, but their measurement remains time-consuming and requires a noise-free environment. Although machine learning has been shown to reduce the required measurement points and, thus, the measurement time, a controlled environment is still necessary. This paper proposes a method to address this constraint by presenting a novel technique that can upsample sparse, noisy HRTF measurements. The proposed approach combines an HRTF Denoisy U-Net for denoising and an Autoencoding Generative Adversarial Network (AE-GAN) for upsampling from three measurement points. The proposed method achieves a log-spectral distortion (LSD) error of 5.41 dB and a cosine similarity loss of 0.0070, demonstrating the method's effectiveness in HRTF upsampling.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。