用视觉语言模型融合摄像头与传感器数据,实时检测自动驾驶中的卫星欺骗攻击。
Development of Vision-Language Model-based GNSS Spoofing Detection for Autonomous Vehicle Navigation

- 通过三阶段微调,将视觉和传感器数据对齐到统一语义空间,识别定位异常
- 在真实道路数据上,攻击检测F1达94%-95%,远超零样本基线(23%-32%)
- 可降低90%计算开销,适合部署在车载系统中作为实时防御层
自动驾驶依赖全球导航卫星系统(GNSS)进行定位与导航,易受欺骗攻击影响,导致车辆被隐蔽引导或引发危险操作。本文首次提出基于视觉语言模型(VLM)的车载欺骗检测框架,融合前视摄像头图像与车内传感器数据(如速度、加速度、航向角),对比GNSS推导轨迹与多源感知数据的一致性。方法采用三阶段微调流程:先锚定视觉线索,再在共享语义空间校准传感器数据,以检测三类攻击场景下的轨迹偏差。研究团队在阿拉巴马州图斯卡卢萨市公共道路实测采集了同步的GNSS、IMU与相机数据,构建独立真实世界验证集,用于评估跨区域泛化能力。在此基础上生成智能欺骗攻击:包括道路网络对齐的镜像轨迹(误转弯)、位置冻结(超速)及漂移生成(停车)。结果显示,零样本VLM基线F1为23%-32%,而本方法达到94%-95%;对误转弯与停车攻击分类准确率100%,超速攻击准确率达88%-93%。此外引入自适应推理策略,将VLM调用频率降至14%(约节省86%计算量),每4秒窗口延迟仅65ms-73ms。结果表明,该方法可作为车载实用级防御层,补充信号级完整性校验。
原文摘要 · Abstract (English)
Autonomous vehicles (AVs) depend on Global Navigation Satellite Systems (GNSS) for localization and navigation, making them vulnerable to spoofing attacks that can covertly redirect vehicles or induce unsafe maneuvers. In this paper, we develop the first Vision-Language Model (VLM)-based framework for GNSS spoofing detection for autonomous vehicles by fusing front-camera visual data with in-vehicle sensor readings (e.g., speed, acceleration, yaw rate) against GNSS-derived maneuvers. Our approach introduces a three-stage fine-tuning process that first grounds visual cues, and then calibrates sensor data within a shared semantic space to detect discrepancies between predicted and GNSS-derived maneuvers across three attack scenarios. We also generated an independent real-world dataset by driving an instrumented vehicle on public roads in Tuscaloosa, Alabama, equipped with time-synchronized GNSS, IMU, and camera logs to validate cross-regional generalization of our fine-tuned model on unseen data from training data. On this dataset, we then generated intelligent spoofing attacks, including trajectory mirroring with road-network snapping for wrong-turn attacks, position freezing for overshoot scenarios, and drift generation for stop attacks. On this validation dataset, the zero-shot VLMs baseline F1-score ranges from 23% to 32%, whereas our fine-tuned model achieves an F1-score ranging from 94% to 95%. Results show that our VLM-based approach correctly classified every wrong-turn and stop attacks, and attains 88%-93% accuracy for overshoot attacks. Furthermore, we introduce an adaptive inference policy that reduces VLM invocations to 14% (~86% computational reduction) and yields 65ms-73ms per 4s window. These results point to a practical, on-road layer of defense that complements signal-level integrity checks with the use of VLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。