EigeNet通过几何引导的多模态学习,实现少样本新视角声学响应精准预测。
EigeNet: Geometry-Informed Multi-Modal Learning for Few-shot Novel View RIR Prediction

- 采用跨视图交替注意力变压器,融合局部声学结构与全局空间关系。
- 在模拟与真实数据上均达当前最优,少样本下仍保持高精度。
- 适合做沉浸式音频渲染、声学建模的研究者与工程师。
从稀疏观测中预测空间变化的房间脉冲响应(RIR)是沉浸式空间音频渲染中的关键但极具挑战性的逆问题。本文提出EIGENET,一种几何引导的多模态少样本新视角RIR预测框架。核心为跨视图交替注意力变压器,可迭代优化局部视图内声学结构与全局跨视图空间关系。受声学射线追踪启发,设计几何引导调制模块,建立几何特征与RIR功率谱间的关联。同时引入辅助损失,将单目标波形预测转化为多任务学习框架。消融实验表明,该设计在不同主干网络下均带来稳定性能提升,验证其基础有效性与架构无关的通用性。在模拟与真实世界基准上评估,EIGENET在少样本新视角RIR预测中达到当前最优,并具备良好的仿真到现实泛化能力。代码与模型权重已公开于https://github.com/FEAfeatherTHER/EigeNet。
原文摘要 · Abstract (English)
Predicting spatially varying Room Impulse Response (RIR) from sparse observations is a critical but highly challenging inverse problem for immersive spatial audio rendering. In this work, we present EIGENET, a geometry-informed multi-modal framework for few-shot novel view RIR prediction. At its core is a Cross-view Alternate-attention Transformer that iteratively refines local intra-view acoustic structures and global cross-view spatial relationships. We empirically demonstrate that this architecture is capable of making full use of the multi-view multi-modal context while performing spatial-temporal reasoning for RIR prediction. Inspired by acoustic ray tracing, we design a geometry-informed modulation block to formulate the connection between geometric features and RIR power spectrum. In the mean time, an auxiliary loss is introduced to transform the single-target waveform prediction into a multi-task learning framework. Through ablation studies, we demonstrate that this design yields consistent performance gains regardless of the underlying backbone, thereby confirming its foundational utility and architecture-agnostic generalizability for RIR prediction task. Evaluated on both simulated and real-world benchmarks, EIGENET achieves both state-of-the-art performance in few-shot novel view RIR prediction and sim-to-real generalization. Codes and checkpoints are available on https://github.com/FEAfeatherTHER/EigeNet.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。