arXiv:2512.10376cs.CV2025-12AAAI被引 2

首个融合4D雷达与激光雷达的场景流估计方法,提升动态环境感知能力。

RaLiFlow: Scene Flow Estimation with 4D Radar and LiDAR Point Clouds

  • 设计动态感知的双向跨模态融合模块,实现雷达与激光雷达信息互补。
  • 构建首个雷达-激光雷达联合场景流数据集,提升动态目标运动估计精度。
  • 适用于自动驾驶中复杂天气下的高鲁棒性场景流分析,适合研发人员参考。

近年来,融合图像与激光雷达点云的多模态方法在场景流估计中展现出潜力。然而,4D毫米波雷达与激光雷达的融合尚未被探索。相比激光雷达,雷达成本更低、抗恶劣天气能力强,且可检测点级速度,是激光雷达的重要补充。但雷达输入存在噪声大、分辨率低、稀疏等问题。此外,目前尚无专门用于场景流估计的雷达-激光雷达数据集。为此,我们基于公开的真实车载数据集构建了雷达-激光雷达场景流数据集。提出有效的雷达去噪预处理策略及场景流标签生成方法,从物体边界推导出更可靠的雷达点流真值。同时,提出RaLiFlow,首个针对4D雷达与激光雷达的联合场景流学习框架,通过创新的动态感知双向跨模态融合(DBCF)模块和精心设计的损失函数,实现有效融合。DBCF模块将雷达动态信息引入局部交叉注意力机制,促进跨模态上下文传播。损失函数缓解训练中不可靠雷达数据的影响,并增强两模态在实例层面的场景流一致性,尤其在动态前景区域表现优异。在重构后的场景流数据集上进行的大量实验表明,该方法显著优于现有的激光雷达和雷达单模态方法。

原文摘要 · Abstract (English)

Recent multimodal fusion methods, integrating images with LiDAR point clouds, have shown promise in scene flow estimation. However, the fusion of 4D millimeter wave radar and LiDAR remains unexplored. Unlike LiDAR, radar is cheaper, more robust in various weather conditions and can detect point-wise velocity, making it a valuable complement to LiDAR. However, radar inputs pose challenges due to noise, low resolution, and sparsity. Moreover, there is currently no dataset that combines LiDAR and radar data specifically for scene flow estimation. To address this gap, we construct a Radar-LiDAR scene flow dataset based on a public real-world automotive dataset. We propose an effective preprocessing strategy for radar denoising and scene flow label generation, deriving more reliable flow ground truth for radar points out of the object boundaries. Additionally, we introduce RaLiFlow, the first joint scene flow learning framework for 4D radar and LiDAR, which achieves effective radar-LiDAR fusion through a novel Dynamic-aware Bidirectional Cross-modal Fusion (DBCF) module and a carefully designed set of loss functions. The DBCF module integrates dynamic cues from radar into the local cross-attention mechanism, enabling the propagation of contextual information across modalities. Meanwhile, the proposed loss functions mitigate the adverse effects of unreliable radar data during training and enhance the instance-level consistency in scene flow predictions from both modalities, particularly for dynamic foreground areas. Extensive experiments on the repurposed scene flow dataset demonstrate that our method outperforms existing LiDAR-based and radar-based single-modal methods by a significant margin.

场景流估计多模态融合自动驾驶4D雷达

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。