用视觉惯性数据训练毫米波雷达场景流,无需昂贵激光雷达。
VISC: mmWave Radar Scene Flow Estimation using Pervasive Visual-Inertial Supervision
- 用视觉惯性融合运动模型与神经网络,消除系统漂移。
- 在烟雾环境中的表现超越依赖激光雷达的顶尖方法。
- 适合智能汽车大规模部署的低成本场景流估计。
本文提出一种基于广泛可用的视觉惯性(VI)传感器数据监督的毫米波雷达场景流估计框架,支持智能汽车的众包训练数据。当前毫米波雷达场景流估计多依赖昂贵且稀缺的3D激光雷达密集点云,而视觉图像难以捕捉物体三维运动,导致动态点监督困难;同时,视觉惯性系统的时序漂移会恶化静态点的场景流估计。为此,我们设计了一种无漂移的刚体变换估计算法,融合基于运动模型的自车运动与神经网络学习结果,为雷达刚体变换提供强监督信号,并推断静态点的场景流。进一步构建了光-毫米波联合监督提取模块,通过光学与毫米波测量的联合约束,学习动态点的场景流。大量实验表明,在烟雾环境中,本方法性能甚至优于依赖昂贵激光雷达的前沿方法。
原文摘要 · Abstract (English)
This work proposes a mmWave radar's scene flow estimation framework supervised by data from a widespread visual-inertial (VI) sensor suite, allowing crowdsourced training data from smart vehicles. Current scene flow estimation methods for mmWave radar are typically supervised by dense point clouds from 3D LiDARs, which are expensive and not widely available in smart vehicles. While VI data are more accessible, visual images alone cannot capture the 3D motions of moving objects, making it difficult to supervise their scene flow. Moreover, the temporal drift of VI rigid transformation also degenerates the scene flow estimation of static points. To address these challenges, we propose a drift-free rigid transformation estimator that fuses kinematic model-based ego-motions with neural network-learned results. It provides strong supervision signals to radar-based rigid transformation and infers the scene flow of static points. Then, we develop an optical-mmWave supervision extraction module that extracts the supervision signals of radar rigid transformation and scene flow. It strengthens the supervision by learning the scene flow of dynamic points with the joint constraints of optical and mmWave radar measurements. Extensive experiments demonstrate that, in smoke-filled environments, our method even outperforms state-of-the-art (SOTA) approaches using costly LiDARs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。