arXiv:2608.01535cs.CVcs.RO2026-08

用汽车雷达监督视觉语言模型,实现动态场景的精准运动与速度估计。

STAR-VLM: Spatiotemporal Grounding Vision-Language Models for Motion and Velocity Estimation via Automotive Radar Supervision

论文配图:STAR-VLM: Spatiotemporal Grounding Vision-Language Models for Motion and Velocity Estimation via Automotive Radar Supervision
图 1 · 摘自论文原文
  • 利用雷达的测距和多普勒信息作为无标签真值训练模型。
  • 在运动分类和速度估计任务上均超越现有方法,达到最先进水平。
  • 低成本、易部署,适合真实自动驾驶场景的度量感知建模。

视觉语言模型(VLM)正成为具身智能的关键组件,广泛应用于自动标注和端到端自动驾驶。然而,现有提升时空推理能力的方法常依赖复杂的预处理流程、昂贵的人工标注或合成数据,限制了可扩展性并引入仿真到现实的差距。尽管这些方法提升了时空理解,仍缺乏对动态场景中物体运动的度量推理能力,如以真实单位估算速度。已有工作探索基于激光雷达的度量深度监督以增强空间感知,但未直接解决时间推理问题。本文提出STAR-VLM,一种由汽车雷达监督的框架,通过范围和多普勒测量提供互补的时空监督,增强视觉语言模型的运动推理与度量速度估计能力。汽车雷达成本低、部署广,其测量结果可作为训练中的无标签真值。在驾驶场景实验中,STAR-VLM在运动分类与度量速度估计任务上均达到当前最优性能,甚至优于为各自任务专门设计的方法。结果表明,汽车雷达是构建面向真实自动驾驶的度量感知时空视觉语言模型的可扩展、低成本监督源。

原文摘要 · Abstract (English)

Vision-language models (VLMs) are emerging as a key component of embodied intelligence, with growing applications in auto-labeling and end-to-end autonomous driving. However, existing approaches for improving spatiotemporal reasoning in VLMs often rely on complex preprocessing pipelines, expensive human annotations, or synthetic data, which limit scalability and introduce potential sim-to-real gaps. Moreover, although these methods have improved spatiotemporal understanding, they still lack strong metric reasoning capabilities for dynamic scenes, such as estimating object motion in real-world units. Prior work has explored LiDAR-based metric depth supervision to enhance spatial perception, but it does not directly address temporal reasoning. We introduce STAR-VLM, an automotive radar-supervised framework that enhances spatiotemporal VLMs with motion reasoning and metric velocity estimation for autonomous driving. Automotive radar is a low-cost and widely deployed sensor that provides complementary spatiotemporal supervision through range and Doppler measurements. By leveraging these measurements as label-free ground truth during training, STAR-VLM improves the metric spatiotemporal reasoning ability of VLMs. Through experiments on driving scenarios, we show that STAR-VLM achieves state-of-the-art performance on both motion classification and metric velocity estimation, outperforming even task-specific methods designed for each task. These results highlight automotive radar as a scalable and cost-effective source of supervision for building metric-aware spatiotemporal VLMs for real-world autonomous driving.

视觉语言模型雷达监督运动估计自动驾驶

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。