用视觉语言模型提升自动驾驶中行人行为与场景理解能力
Application of Vision-Language Model to Pedestrians Behavior and Scene Understanding in Autonomous Driving
- 将大模型知识蒸馏到轻量网络,实现高效部署
- 在开放词汇感知和轨迹预测任务上显著提升性能
- 适合关注自动驾驶感知与决策的工程师与研究者
视觉语言模型(VLMs)在增强自动驾驶中的感知与决策方面展现出巨大潜力,但在理解行人交互复杂场景及高效车辆部署方面仍存在差距。本文提出一种知识蒸馏方法,将大规模视觉语言基础模型的知识迁移至高效视觉网络,并应用于行人行为预测与场景理解任务,生成更丰富多样的语义属性。通过融合多个预训练模型并采用集成技术进一步提升性能。经知识蒸馏后,模型在开放词汇感知与轨迹预测任务中表现显著提升,有望改善自动驾驶端到端系统的整体表现。
原文摘要 · Abstract (English)
Vision-language models (VLMs) have become a promising approach to enhancing perception and decision-making in autonomous driving. The gap remains in applying VLMs to understand complex scenarios interacting with pedestrians and efficient vehicle deployment. In this paper, we propose a knowledge distillation method that transfers knowledge from large-scale vision-language foundation models to efficient vision networks, and we apply it to pedestrian behavior prediction and scene understanding tasks, achieving promising results in generating more diverse and comprehensive semantic attributes. We also utilize multiple pre-trained models and ensemble techniques to boost the model's performance. We further examined the effectiveness of the model after knowledge distillation; the results show significant metric improvements in open-vocabulary perception and trajectory prediction tasks, which can potentially enhance the end-to-end performance of autonomous driving.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。