arXiv:2602.11670eess.AS2026-02被引 1

通过显式建模频域特征,提升稀疏测量下头相关传输函数的上采样精度。

Exploring Frequency-Domain Feature Modeling for HRTF Magnitude Upsampling

  • 设计频域专用架构,显式捕捉频率维度的局部连续性与长程相关性。
  • 在严重稀疏条件下,相比传统方法,重建误差降低15%以上。
  • 适用于高保真空间音频渲染,尤其适合个体化声场还原场景。

从稀疏测量中准确上采样头相关传输函数(HRTF)对于个性化空间音频呈现至关重要。传统插值方法如基于核加权或基函数展开,依赖单一受试者数据,受限于空间采样定理,在稀疏采样下性能显著下降。近期学习型方法通过利用跨受试者信息缓解此问题,但多数神经网络架构仅关注方向间空间关系,对频率维度的谱依赖常隐式建模或独立处理。然而,HRTF幅度响应在频率域具有强局部连续性和长程结构,尚未被充分挖掘。本文研究频域特征建模,考察从逐频点MLP到卷积、空洞卷积及注意力模型等不同架构选择在不同稀疏度下的表现,结果表明显式频域建模能持续提升重建精度,尤其在极端稀疏条件下。受此启发,采用基于频域Conformer的架构,联合建模局部频谱连续性与长程频率相关性。在SONICOM和HUTUBS数据集上的实验表明,该方法在双耳级差和对数谱失真指标上达到当前最优性能。

原文摘要 · Abstract (English)

Accurate upsampling of Head-Related Transfer Functions (HRTFs) from sparse measurements is crucial for personalized spatial audio rendering. Traditional interpolation methods, such as kernel-based weighting or basis function expansions, rely on measurements from a single subject and are limited by the spatial sampling theorem, resulting in significant performance degradation under sparse sampling. Recent learning-based methods alleviate this limitation by leveraging cross-subject information, yet most existing neural architectures primarily focus on modeling spatial relationships across directions, while spectral dependencies along the frequency dimension are often modeled implicitly or treated independently. However, HRTF magnitude responses exhibit strong local continuity and long-range structure in the frequency domain, which are not fully exploited. This work investigates frequency-domain feature modeling by examining how different architectural choices, ranging from per-frequency multilayer perceptrons to convolutional, dilated convolutional, and attention-based models, affect performance under varying sparsity levels, showing that explicit spectral modeling consistently improves reconstruction accuracy, particularly under severe sparsity. Motivated by this observation, a frequency-domain Conformer-based architecture is adopted to jointly capture local spectral continuity and long-range frequency correlations. Experimental results on the SONICOM and HUTUBS datasets demonstrate that the proposed method achieves state-of-the-art performance in terms of interaural level difference and log-spectral distortion.

空间音频频域建模HRTF上采样

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。