用潜在扩散模型生成高分辨率雷达点云,提升恶劣环境感知能力
RaLD: Generating High-Resolution 3D Radar Point Clouds with Latent Diffusion
- 基于雷达频谱直接生成点云,结合激光雷达编码与不变序潜空间
- 生成点云密度和精度显著优于传统方法,保留结构细节
- 适合自动驾驶等需鲁棒3D感知的场景
毫米波雷达因其在恶劣条件下的鲁棒性和低成本,成为自动驾驶系统有前景的感知模态。然而,其应用受限于点云稀疏和低分辨率,难以满足高精度3D感知任务需求。尽管已有研究尝试通过生成方法改善,但多依赖密集体素表示,效率低且难保留结构细节。本文观察到,潜在扩散模型(LDM)虽在其他模态成功,却因缺乏兼容表示与条件策略未被有效用于雷达3D生成。为此提出RaLD框架,融合场景级视锥激光雷达自编码、顺序不变潜表示及直接雷达频谱条件化,实现更紧凑高效的生成过程。实验表明,RaLD能从原始雷达频谱生成稠密准确的3D点云,为复杂环境下的鲁棒感知提供可行方案。
原文摘要 · Abstract (English)
Millimeter-wave radar offers a promising sensing modality for autonomous systems thanks to its robustness in adverse conditions and low cost. However, its utility is significantly limited by the sparsity and low resolution of radar point clouds, which poses challenges for tasks requiring dense and accurate 3D perception. Despite that recent efforts have shown great potential by exploring generative approaches to address this issue, they often rely on dense voxel representations that are inefficient and struggle to preserve structural detail. To fill this gap, we make the key observation that latent diffusion models (LDMs), though successful in other modalities, have not been effectively leveraged for radar-based 3D generation due to a lack of compatible representations and conditioning strategies. We introduce RaLD, a framework that bridges this gap by integrating scene-level frustum-based LiDAR autoencoding, order-invariant latent representations, and direct radar spectrum conditioning. These insights lead to a more compact and expressive generation process. Experiments show that RaLD produces dense and accurate 3D point clouds from raw radar spectrums, offering a promising solution for robust perception in challenging environments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。