arXiv:2608.04130cs.CV2026-08

用雷达点云实现无需相机的4D动态场景理解,突破传统视觉依赖。

Radar4D-VLM: Proposal-Grounded Temporal 4D Radar Reasoning Across Frozen Language Models

论文配图:Radar4D-VLM: Proposal-Grounded Temporal 4D Radar Reasoning Across Frozen Language Models
图 1 · 摘自论文原文
  • 基于连续10帧雷达数据生成物体提议,构建层次化语义-运动令牌
  • 在K-Radar上达到98.13%的物体召回率,优于基线6.4~22.8个百分点
  • 适配五大语言模型家族,验证接口兼容性,不依赖语言监督提升性能

自动驾驶中的视觉-语言模型主要依赖摄像头和激光雷达,而4D雷达虽具备恶劣天气鲁棒性和径向速度直接测量能力,却长期未被作为独立感知模态开发。本文提出Radar4D-VLM,一种仅使用雷达的时序视觉-语言模型,可处理连续十帧4D雷达点云,无需相机或激光雷达输入。该模型提取几何对齐的物体提议,并将雷达证据组织为物体、场景与运动状态三类紧凑令牌。通过参数高效投影器将这些令牌映射至冻结的语言模型主干,同时可解释的预测头联合建模物体数量、空间分布、运动状态、碰撞风险、语义类别及径向速度。Radar4D-VLM在序列隔离的K-Radar开发集上,4米距离下Top-64提案召回率达98.13%,高于固定网格和均匀随机基线6.40和22.83个百分点。我们进一步在八个相同适应预算下评估五类冻结语言模型(Qwen、Phi、Mistral、Llama、Gemma)共24组匹配运行。结果表明,雷达令牌接口在所有五类模型中均保持兼容性,而对照实验显示语言监督并未带来稳定性能增益,证实了接口通用性与语言作用的解耦。该工作建立了雷达单模态多模态推理的可复现基础。

原文摘要 · Abstract (English)

Vision-language models for autonomous driving primarily rely on cameras and LiDAR, leaving 4D radar largely unexplored as a standalone perceptual modality despite its robustness to adverse visibility and direct measurement of radial velocity. We introduce Radar4D-VLM, a radar-only temporal vision-language model that reasons from ten consecutive 4D-radar point-cloud sweeps without camera or LiDAR input. Radar4D-VLM extracts geometrically grounded object proposals and organizes radar evidence into a compact hierarchy of object, scene, and kinematic tokens. A parameter-efficient projector maps these tokens into frozen language backbones, while auditable prediction heads jointly model object count, spatial distribution, motion state, collision risk, semantic category, and radial velocity. Radar4D-VLM combines proposal-grounded temporal object tokenization, global scene context, and explicit kinematic tokens within a unified frozen-backbone interface. On sequence-isolated K-Radar development validation, its Top-64 proposal recall reaches 98.13% at 4 m, exceeding fixed-lattice and uniform-random controls by 6.40 and 22.83 percentage points, respectively. We further evaluate 24 matched runs spanning eight frozen Qwen, Phi, Mistral, Llama, and Gemma backbones under an identical adaptation budget. The radar-token interface remains compatible across all five language-model families, while matched aligned, permuted, and no-language controls show sensor dependence but no stable direct-head gain from aligned language supervision. These results establish a reproducible foundation for radar-only multimodal scene and motion reasoning while separating interface compatibility from the benefit of language supervision.

4D雷达视觉语言模型自动驾驶时序推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。