神经网络在时差定位中超越传统方法,关键在于自适应频率加权。
What Do Neural Networks Learn for TDOA Estimation? A Cross-Architecture Probing Study

- 通过探测网络内部机制,发现其学习了保留可靠性信息的频率加权策略。
- 移除PHAT步骤后,无论是经典还是神经网络方法性能均提升,尤其在噪声环境下。
- 适合研究语音增强、声源定位或模型可解释性的研究人员参考。
神经网络在噪声和混响环境下对时差到达(TDOA)估计的表现优于传统GCC-PHAT方法,但其内部学习机制尚不明确。为揭示这一机制,我们将GCC-PHAT的数学步骤转化为诊断目标,对MLP、CNN和Transformer三种架构的隐藏层进行探测,并结合梯度归因与因果频率掩码分析。结果表明,跨频谱功率计算在所有架构和条件下均稳定出现,而定义GCC-PHAT的核心步骤——PHAT白化却未显现。相反,网络学习了一种考虑幅度信息的频率加权方式,保留了被PHAT丢弃的每频段可靠性信息。这表明PHAT是一个信息瓶颈:无论在经典还是神经网络的GCC流程中移除该步骤,都能在加性噪声下提升性能。在真实混响数据上,PHAT仍是经典方法中的最优加权方式,但端到端网络通过学习数据自适应加权实现了更低的误差。
原文摘要 · Abstract (English)
Neural networks outperform classical GCC-PHAT for Time-Difference-of-Arrival (TDOA) estimation in noise and reverberation, yet their internal strategy remains unexplored. To uncover it, we turn GCC-PHAT's mathematical steps into diagnostic targets, probing hidden layers of three architectures (MLP, CNN, Transformer) and complementing with gradient attribution and causal frequency masking. We find that cross-power computation consistently emerges across all architectures and conditions, while PHAT whitening, the defining step of GCC-PHAT, fails to emerge. Instead, networks learn a magnitude-aware frequency weighting that preserves per-frequency reliability information discarded by PHAT. This makes PHAT an information bottleneck: removing it from both classical and neural GCC pipelines improves performance under additive noise. On real-world reverberant data, PHAT remains the best classical weighting, but end-to-end networks achieve lower error by learning data-adaptive weighting.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。