arXiv:2602.13314cs.CVcs.AI2026-02

用视觉语言模型生成雷达数据,解决真实雷达数据少的问题。

Sim2Radar: Toward Bridging the Radar Sim-to-Real Gap with VLM-Guided Scene Reconstruction

  • 从单张图像重建材质感知3D场景,再用物理模型模拟毫米波信号。
  • 在真实室内场景上训练检测模型,3D AP提升最高达+3.7(IoU 0.3)。
  • 适合做雷达感知但缺乏标注数据的研究者,尤其关注跨域迁移。

毫米波雷达在烟雾、尘埃和低光等视觉退化环境中仍能可靠感知,但基于学习的雷达感知受限于大规模雷达数据的稀缺与标注成本。本文提出Sim2Radar,一种端到端框架,可直接从单视角RGB图像合成训练用雷达数据,实现无需人工建模的可扩展数据生成。该方法结合单目深度估计、分割与视觉-语言推理,推断物体材质,进而利用可配置的物理基射线追踪器,基于ITU-R电磁属性参数化的Fresnel反射模型模拟毫米波传播。在真实室内场景上的评估表明,通过迁移学习:先在合成数据上预训练雷达点云目标检测模型,再在真实雷达数据上微调,3D AP(IoU 0.3)最高提升3.7,增益主要来自空间定位性能改善。结果表明,基于物理、由视觉驱动的雷达仿真可为雷达学习提供有效的几何先验,并在真实数据有限时显著提升性能。

原文摘要 · Abstract (English)

Millimeter-wave (mmWave) radar provides reliable perception in visually degraded indoor environments (e.g., smoke, dust, and low light), but learning-based radar perception is bottlenecked by the scarcity and cost of collecting and annotating large-scale radar datasets. We present Sim2Radar, an end-to-end framework that synthesizes training radar data directly from single-view RGB images, enabling scalable data generation without manual scene modeling. Sim2Radar reconstructs a material-aware 3D scene by combining monocular depth estimation, segmentation, and vision-language reasoning to infer object materials, then simulates mmWave propagation with a configurable physics-based ray tracer using Fresnel reflection models parameterized by ITU-R electromagnetic properties. Evaluated on real-world indoor scenes, Sim2Radar improves downstream 3D radar perception via transfer learning: pre-training a radar point-cloud object detection model on synthetic data and fine-tuning on real radar yields up to +3.7 3D AP (IoU 0.3), with gains driven primarily by improved spatial localization. These results suggest that physics-based, vision-driven radar simulation can provide effective geometric priors for radar learning and measurably improve performance under limited real-data supervision.

雷达感知仿真生成跨域迁移视觉语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。