arXiv:2409.14935cs.CV2024-09ECCV被引 1

用注意力机制融合多视角深度残缺视频,提升补全精度。

Deep Cost Ray Fusion for Sparse Depth Video Completion

论文配图:Deep Cost Ray Fusion for Sparse Depth Video Completion
图 1 · 摘自论文原文
  • 构建深度假设平面的成本体积,通过光线对齐融合多视角信息。
  • 在KITTI、VOID、ScanNetV2上超越或媲美顶尖方法,参数更少。
  • 适合需要高效高精度深度补全的自动驾驶与三维重建场景。

本文提出一种基于学习的稀疏深度视频补全框架。给定特定视角下的稀疏深度图和彩色图像,我们的方法构建基于深度假设平面的成本体积。为有效融合多个视角的时序成本体积以提升深度补全效果,我们引入一种名为RayFusion的学习型成本体积融合框架,该框架针对相邻成本体积中重叠光线对,利用注意力机制进行特征融合。得益于随时间累积的特征统计信息,所提框架在多样化的室内与室外数据集上(包括KITTI Depth Completion、VOID Depth Completion及ScanNetV2)持续优于或媲美现有最先进方法,且使用更少网络参数。

原文摘要 · Abstract (English)

In this paper, we present a learning-based framework for sparse depth video completion. Given a sparse depth map and a color image at a certain viewpoint, our approach makes a cost volume that is constructed on depth hypothesis planes. To effectively fuse sequential cost volumes of the multiple viewpoints for improved depth completion, we introduce a learning-based cost volume fusion framework, namely RayFusion, that effectively leverages the attention mechanism for each pair of overlapped rays in adjacent cost volumes. As a result of leveraging feature statistics accumulated over time, our proposed framework consistently outperforms or rivals state-of-the-art approaches on diverse indoor and outdoor datasets, including the KITTI Depth Completion benchmark, VOID Depth Completion benchmark, and ScanNetV2 dataset, using much fewer network parameters.

深度补全多视角融合注意力机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。