用4D毫米波雷达实现雨雪天气下的稳定场景语义理解
Weather-Robust Scene Semantics with Vision-Aligned 4D Radar

- 将雷达编码器对齐视觉预训练嵌入,通过冻结的视觉语言模型生成场景描述
- 在雾、小雪、大雪条件下,雷达方案远超摄像头(幻觉率超90%)
- 发现并解决雷达与视觉模型间特征归一化不匹配问题,适合自动驾驶感知研究
相机和激光雷达在雨、雾、雪中性能下降,而毫米波雷达则基本不受影响。本文将雷达编码器对齐冻结的SigLIP视觉嵌入,并通过仅约700万可训练参数的冻结视觉语言模型(VLM)解码出结构化场景描述。在包含未见雾、轻雪和重雪序列的K-RADAR数据集上,所有雷达配置均显著优于摄像头基线(幻觉率超过90%)。我们识别出跨模态对齐中的令牌归一化不匹配是主要失败原因,并证明投影输出层归一化可有效解决。对编码器复杂度、描述格式和池化策略的分析揭示了未来雷达-VLM流水线设计的关键权衡。
原文摘要 · Abstract (English)
Cameras and LiDAR degrade in rain, fog, and snow, while millimeter-wave radar remains largely unaffected. We align a radar encoder to frozen SigLIP vision embeddings and decode structured scene captions through a frozen vision-language model (VLM) with approximately 7M trainable parameters. On K-RADAR with held-out fog, light snow, and heavy snow sequences, all radar configurations outperform a camera baseline that collapses to over 90% hallucination. We identify a token-norm mismatch as the dominant failure mode when bridging radar to a frozen VLM and show that projector-output LayerNorm resolves it. Analysis of encoder complexity, caption format, and pooling strategy reveals tradeoffs that inform future radar-VLM pipeline design.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。