融合点与体素信息,提升3D场景光流估计精度
Self-Supervised Scene Flow Estimation with Point-Voxel Fusion and Surface Representation
- 设计点-体素双分支结构,兼顾长程依赖与细节特征
- 在KITTI数据集上EPE指标降低超8.5%,性能超越现有自监督方法
- 适合需要高精度3D运动分析的自动驾驶、机器人领域
场景光流估计旨在生成连续两帧点云间的3D运动场,应用广泛。现有基于点的方法忽略点云不规则性,难以捕捉长程依赖;基于体素的方法则易丢失细节信息。本文提出一种点-体素融合方法:利用基于稀疏网格注意力和移位窗口策略的体素分支捕捉长程依赖,点分支提取细粒度特征以弥补体素分支的信息损失。此外,为更好描述复杂3D物体的几何结构,引入伞状表面特征提取(USFE)模块显式编码局部表面信息。在Flyingthings3D和KITTI数据集上的实验表明,本方法优于所有自监督方法,在完全监督方法中也表现优异。各项指标均有提升,尤其在KITTIo和KITTIs数据集上EPE分别降低8.51%和10.52%。
原文摘要 · Abstract (English)
Scene flow estimation aims to generate the 3D motion field of points between two consecutive frames of point clouds, which has wide applications in various fields. Existing point-based methods ignore the irregularity of point clouds and have difficulty capturing long-range dependencies due to the inefficiency of point-level computation. Voxel-based methods suffer from the loss of detail information. In this paper, we propose a point-voxel fusion method, where we utilize a voxel branch based on sparse grid attention and the shifted window strategy to capture long-range dependencies and a point branch to capture fine-grained features to compensate for the information loss in the voxel branch. In addition, since xyz coordinates are difficult to describe the geometric structure of complex 3D objects in the scene, we explicitly encode the local surface information of the point cloud through the umbrella surface feature extraction (USFE) module. We verify the effectiveness of our method by conducting experiments on the Flyingthings3D and KITTI datasets. Our method outperforms all other self-supervised methods and achieves highly competitive results compared to fully supervised methods. We achieve improvements in all metrics, especially EPE, which is reduced by 8.51% on the KITTIo dataset and 10.52% on the KITTIs dataset, respectively.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。