arXiv:2501.13017eess.AScs.SD2025-01中稿 · ICASSP 2025被引 10

用检索增强神经场,仅凭3~5个测量点就能高精度还原个性化头相关传输函数。

Retrieval-Augmented Neural Field for HRTF Upsampling and Personalization

  • 从数据库中检索相似受试者,将其HRTF作为额外输入增强神经场。
  • 在SONICOM数据集上,仅用3个测量方向即实现接近全网格的还原效果。
  • 适合需要低采样成本、高个性化沉浸音效的音频生成与虚拟现实应用。

密集空间网格的头相关传输函数(HRTF)对沉浸式双耳音频生成至关重要,但其录制耗时。尽管神经场在空间上采样方面取得进展,但从少数测量方向(如3或5个)进行空间上采样仍具挑战。为此,我们提出检索增强神经场(RANF)。RANF从数据集中检索与目标受试者HRTF相近的个体,将该个体在目标方向的HRTF输入神经场,同时保留声源方向信息。此外,我们设计了一种受多通道处理技术启发的神经网络,可高效处理多个检索到的受试者。实验表明,RANF在SONICOM数据集上表现优异,且是2024年听觉个性化挑战赛任务2获胜方案的核心组件。

原文摘要 · Abstract (English)

Head-related transfer functions (HRTFs) with dense spatial grids are desired for immersive binaural audio generation, but their recording is time-consuming. Although HRTF spatial upsampling has shown remarkable progress with neural fields, spatial upsampling only from a few measured directions, e.g., 3 or 5 measurements, is still challenging. To tackle this problem, we propose a retrieval-augmented neural field (RANF). RANF retrieves a subject whose HRTFs are close to those of the target subject from a dataset. The HRTF of the retrieved subject at the desired direction is fed into the neural field in addition to the sound source direction itself. Furthermore, we present a neural network that can efficiently handle multiple retrieved subjects, inspired by a multi-channel processing technique called transform-average-concatenate. Our experiments confirm the benefits of RANF on the SONICOM dataset, and it is a key component in the winning solution of Task 2 of the listener acoustic personalization challenge 2024.

音频生成神经场个性化检索增强

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。