四鱼眼相机实时生成360°深度图,速度达20帧/秒。
FastViDAR: Real-Time Omnidirectional Depth Estimation via Alternative Hierarchical Attention
- 用分层注意力机制融合多视角特征,降低计算开销。
- 将多视角深度投影到等距柱状图,实现高质量融合深度。
- 在真实数据上零样本表现优异,适合嵌入式实时应用。
本文提出FastViDAR,一种基于四个鱼眼相机输入的全新框架,可生成完整的$360^ ext{∘}$深度图,以及每台相机的深度、融合深度和置信度估计。主要贡献包括:(1) 提出替代分层注意力(AHA)机制,通过独立的帧内与帧间窗口自注意力,高效融合跨视角特征,实现特征混合的同时降低计算开销;(2) 提出一种新颖的等距柱状图(ERP)融合方法,将多视角深度估计投影至统一的等距柱状坐标系,获得最终融合深度;(3) 利用HM3D和2D3D-S数据集构建等距柱状图图像-深度对,进行全面评估,在真实数据集上展现出具有竞争力的零样本性能,并在NVIDIA Orin NX嵌入式硬件上达到最高20 FPS的处理速度。
原文摘要 · Abstract (English)
In this paper we propose FastViDAR, a novel framework that takes four fisheye camera inputs and produces a full $360^\circ$ depth map along with per-camera depth, fusion depth, and confidence estimates. Our main contributions are: (1) We introduce Alternative Hierarchical Attention (AHA) mechanism that efficiently fuses features across views through separate intra-frame and inter-frame windowed self-attention, achieving cross-view feature mixing with reduced overhead. (2) We propose a novel ERP fusion approach that projects multi-view depth estimates to a shared equirectangular coordinate system to obtain the final fusion depth. (3) We generate ERP image-depth pairs using HM3D and 2D3D-S datasets for comprehensive evaluation, demonstrating competitive zero-shot performance on real datasets while achieving up to 20 FPS on NVIDIA Orin NX embedded hardware. Project page: \href{https://3f7dfc.github.io/FastVidar/}{https://3f7dfc.github.io/FastVidar/}
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。