用大模型动态融合视觉与雷达数据,提升自动驾驶在恶劣条件下的感知鲁棒性。
LLM-Empowered Multimodal Fusion Framework for Autonomous Driving: Semantic Enhancement and Channel-Adaptive Design

- 以大模型为核心,根据信道质量动态调整雷达数据的融合权重。
- 在nuScenes上定位误差降低40%,VIRAT上定位精度达0.214m。
- 适合需要高可靠性多模态融合的自动驾驶系统研发者。
视觉-雷达融合是实现稳健自动驾驶的核心,结合了视觉的丰富语义与雷达的精确距离和速度测量。然而,真实场景中因遮挡、恶劣天气和信道噪声导致输入质量动态变化,严重制约融合效果。为此,我们提出一种以大语言模型(LLM)为中心的语义层信道感知集成感知框架(LM-SCIP),将问题从静态数据融合重构为信道感知的语义推理。该框架将一个分层的雷达-视觉编码器与信道自适应语义模块(CASM)耦合,通过链路指标生成“信道提示”以动态调控外部雷达特征。采用参数高效微调的LoRA-LMM与异构专家混合(H-MoE)协同决策,权衡本地视觉线索与信道调节后的雷达上下文。最后由解耦的多任务解码器输出定位、轨迹预测与图像重建。在nuScenes和VIRAT数据集上的实验验证了方法的有效性:在nuScenes上,雷达输入可控切换时,定位均方根误差(RMSE)相比纯视觉基线降低40.0%;在VIRAT上,实现0.214m定位RMSE与0.179m最小最终位移误差(minFDE, k=1)。结果表明,所提框架可在低信噪比(SNR)下实现可靠的视觉主导容错,在高SNR下达成协同融合。
原文摘要 · Abstract (English)
Vision-radar fusion is central to robust autonomous driving, combining dense visual semantics with precise range and velocity measurements from radar. However, real-world fusion quality is fundamentally challenged by dynamically varying input quality, stemming from occlusion, adverse weather, and channel noise. To address this, we re-frame the problem from static data fusion to channel-aware semantic reasoning and propose a Large Language Model-centric Semantic-layer Channel-aware Integrated Perception (LM-SCIP) framework. It places a Large Language Model (LLM) as a central reasoning core to fuse a local visual stream with a quality-varying external radar stream used to cover perception-blind spots. Concretely, LM-SCIP couples a hierarchical radar-vision encoder with a Channel-Adaptive Semantic Module (CASM) that maps link indicators into a "Channel Prompt" to dynamically gate external radar features. A parameter-efficient, LoRA-tuned LLM, in conjunction with a heterogeneous Mixture-of-Experts (H-MoE), then arbitrates between local visual cues and the channel-conditioned radar context. Finally, a decoupled multi-task decoder outputs localization, trajectory forecasting, and image reconstruction. Experiments on nuScenes and VIRAT validate our approach. On nuScenes, under a controlled toggle of radar input, LM-SCIP reduces localization RMSE by 40.0% versus a vision-only baseline. On VIRAT, the model attains a 0.214m localization RMSE and 0.179m minFDE (k=1). These results reveal that the proposed LM-SCIP enables a robust vision-dominant fallback at low SNR and synergistic fusion at high SNR.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。