arXiv:2602.01861eess.AScs.LG2026-02被引 3

用Transformer实现任意麦克风位置的声学响应重建

RIR-Former: Coordinate-Guided Transformer for Continuous Reconstruction of Room Impulse Responses

  • 基于正弦编码与分段多分支解码器,直接建模麦克风位置信息
  • 在多种缺失率下,NMSE和余弦距离均优于现有方法
  • 适合需要灵活阵列布局的声学系统部署

房间冲激响应(RIR)对诸多声学信号处理任务至关重要,但密集空间测量常不切实际。本文提出RIR-Former,一种无需网格、单步前馈的RIR重建模型。通过在Transformer主干中引入正弦编码模块,有效融合麦克风位置信息,支持任意阵列位置的插值。此外,设计分段多分支解码器,分别处理早期反射与晚期混响,提升全时域重建效果。在多样化的模拟声学环境中实验表明,无论缺失率或阵列配置如何变化,RIR-Former在归一化均方误差(NMSE)和余弦距离(CD)上均持续优于当前最优基线。结果验证了该方法在实际部署中的潜力,并为未来向随机线性阵列、复杂几何结构、动态场景及真实环境扩展提供启发。

原文摘要 · Abstract (English)

Room impulse responses (RIRs) are essential for many acoustic signal processing tasks, yet measuring them densely across space is often impractical. In this work, we propose RIR-Former, a grid-free, one-step feed-forward model for RIR reconstruction. By introducing a sinusoidal encoding module into a transformer backbone, our method effectively incorporates microphone position information, enabling interpolation at arbitrary array locations. Furthermore, a segmented multi-branch decoder is designed to separately handle early reflections and late reverberation, improving reconstruction across the entire RIR. Experiments on diverse simulated acoustic environments demonstrate that RIR-Former consistently outperforms state-of-the-art baselines in terms of normalized mean square error (NMSE) and cosine distance (CD), under varying missing rates and array configurations. These results highlight the potential of our approach for practical deployment and motivate future work on scaling from randomly spaced linear arrays to complex array geometries, dynamic acoustic scenes, and real-world environments.

声学重建Transformer麦克风阵列信号处理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。