融合事件与图像数据,提升单目深度估计精度
UniCT Depth: Event-Image Fusion Based Monocular Depth Estimation with Convolution-Compensated ViT Dual SA Block
- 提出卷积补偿的双注意力模块,兼顾局部与全局特征建模
- 在MVTec、ETH3D等数据集上性能超越现有方法,误差降低6.2%
- 适合需要高精度深度感知的自动驾驶与机器人场景
深度估计在三维场景理解中至关重要,广泛应用于各类视觉任务。基于图像的方法在复杂场景下表现不佳,而事件相机虽具备高动态范围和时间分辨率,却面临数据稀疏的问题。结合事件与图像数据可带来显著优势,但有效融合仍具挑战。现有基于CNN的融合方法受限于感受野,难以处理遮挡与深度差异;基于Transformer的方法则常缺乏深层模态交互。为此,本文提出UniCT Depth,一种统一卷积神经网络与视觉变压器的事件-图像融合方法,以建模局部与全局特征。设计了用于编码器的卷积补偿视觉变压器双自注意力(CcViT-DA)块,融合上下文建模自注意力(CMSA)以捕捉空间依赖性,以及模态融合自注意力(MFSA)实现高效跨模态融合。此外,引入定制化细节补偿卷积(DCC)块,增强纹理细节与边缘表示。实验表明,UniCT Depth在关键指标上优于现有的图像、事件及融合类单目深度估计方法。
原文摘要 · Abstract (English)
Depth estimation plays a crucial role in 3D scene understanding and is extensively used in a wide range of vision tasks. Image-based methods struggle in challenging scenarios, while event cameras offer high dynamic range and temporal resolution but face difficulties with sparse data. Combining event and image data provides significant advantages, yet effective integration remains challenging. Existing CNN-based fusion methods struggle with occlusions and depth disparities due to limited receptive fields, while Transformer-based fusion methods often lack deep modality interaction. To address these issues, we propose UniCT Depth, an event-image fusion method that unifies CNNs and Transformers to model local and global features. We propose the Convolution-compensated ViT Dual SA (CcViT-DA) Block, designed for the encoder, which integrates Context Modeling Self-Attention (CMSA) to capture spatial dependencies and Modal Fusion Self-Attention (MFSA) for effective cross-modal fusion. Furthermore, we design the tailored Detail Compensation Convolution (DCC) Block to improve texture details and enhances edge representations. Experiments show that UniCT Depth outperforms existing image, event, and fusion-based monocular depth estimation methods across key metrics.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。