arXiv:2603.27998eess.AScs.LG2026-03中稿 · Interspeech 2026, …

用稀疏测量重建任意方向头相关脉冲响应,提升听觉还原精度。

HRIR-Former: Grid-Free Time-Domain Reconstruction of Head-Related Impulse Responses with a Spatially Encoded Transformer

  • 基于时域、无网格的Transformer架构,直接预测任意方向声波响应。
  • 在SONICOM数据集上,NMSE和角度误差均优于现有方法。
  • 无需最小相位假设,适合个性化音频渲染场景。

个性化头相关脉冲响应(HRIR)可实现沉浸式双耳渲染,但针对每位用户密集测量成本高昂。本文提出从稀疏测量中进行空间上采样:给定少数已测方向的HRIR,预测未测量方向的响应。以往方法多在频域操作,依赖最小相位假设或独立的时间模型,且使用固定方向网格,影响时域保真度与空间连续性。我们提出HRIR-Former,一种时域、无网格的双耳Transformer,可从稀疏输入重建任意方向的HRIR。其采用正弦空间特征、1D卷积精炼模块,并引入辅助的双耳时间差(ITD)与双耳级差(ILD)预测头。在SONICOM数据集上,该方法在归一化均方误差(NMSE)、余弦距离及ITD/ILD误差方面均优于现有方法;消融实验验证了各模块有效性,表明最小相位预处理并非必需。

原文摘要 · Abstract (English)

Individualized head-related impulse responses (HRIRs) enable binaural rendering, but dense per-listener measurements are costly. We address HRIR spatial up-sampling from sparse per-listener measurements: given a few measured HRIRs for a listener, predict HRIRs at unmeasured target directions. Prior learning methods often work in the frequency domain, rely on minimum-phase assumptions or separate timing models, and use a fixed direction grid, which can degrade temporal fidelity and spatial continuity. We propose HRIR-Former, a time-domain, grid-free binaural Transformer for reconstructing HRIRs at arbitrary directions from sparse inputs. It uses sinusoidal spatial features, a Conv1D refinement module, and auxiliary interaural time difference (ITD) and interaural level difference (ILD) heads. On SONICOM, it improves normalized mean squared error (NMSE), cosine distance, and ITD/ILD errors over prior methods; ablations validate modules and show minimum-phase preprocessing is unnecessary.

音频生成Transformer双耳渲染

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。