arXiv:2511.21105cs.CV2025-11被引 2

用语言描述雷达场景,让模型理解物体空间位置。

RLM: A Vision-Language Model Approach for Radar Scene Understanding

  • 用结构化语言标签描述雷达中车辆分布,实现空间对齐。
  • 新目标函数使雷达图与语言匹配更精确,提升定位能力。
  • 适合做自动驾驶感知、多模态融合的研究者参考。

雷达传感器在恶劣天气、光照和远距离条件下仍能提供可靠感知,但现有机器学习方法分散且任务专用,各下游任务使用不同架构与训练目标。本文提出RadarVLM,一种基于视觉-语言的统一场景理解框架,通过结构化空间语言监督学习统一的场景表征。利用带有真实雷达模型的CARLA模拟器,在多样化场景中模拟超过110小时驾驶,收集了800,000+组雷达-文本配对数据。主要贡献包括:(1)一种将车辆分布编码为雷达原生坐标系的语言框架;(2)空间锚定的CLIP(SG-CLIP)目标函数,以连续场景相似性替代二元匹配,支持细粒度空间推理。此外,提出考虑定位精度的评估指标,超越传统语言相似性度量。在生成式描述与车辆分割任务上验证,SG-CLIP相比基线CLIP相对F1分数提升50%,分割平均精度(AP)提高21%,证明语言引导可生成具有空间结构的表示。

原文摘要 · Abstract (English)

Radar sensors provide reliable perception across adverse weather, lighting, and long-range conditions, yet existing machine learning approaches remain fragmented and task-specific, with each downstream task employing distinct architectures and training objectives. We present RadarVLM, a vision-language framework that learns unified scene-level representations through structured spatial language supervision. Leveraging the CARLA simulator with a realistic radar model, we collect over 800k radar-caption pairs across 110+ hours of simulated driving in diverse scenarios. We make two key contributions: (1) a structured caption framework encoding vehicle distributions in the radar's native coordinate system, and (2) Spatially-Grounded CLIP (SG-CLIP) objective that replaces binary matching with continuous scene similarity, enabling fine-grained spatial reasoning. We further propose localization-aware evaluation metrics that directly assess spatial accuracy beyond traditional linguistic similarity measures. Validated on generative captioning and vehicle segmentation, SG-CLIP achieves up to 50% relative F1-score improvement over vanilla CLIP and a 21% AP gain on segmentation, demonstrating that language grounding produces spatially structured representations.

雷达感知多模态空间推理自动驾驶

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。