arXiv:2506.02014cs.CVcs.AI2025-06

用动态提示优化提升自动驾驶场景理解能力。

Research on Driving Scenario Technology Based on Multimodal Large Lauguage Model Optimization

  • 根据图像内容动态调整提示,聚焦影响车辆的关键物体。
  • 融合真实与合成数据构建高质量训练集,提升模型泛化性。
  • 结合知识蒸馏等技术降低资源消耗,适合落地部署。

随着自动驾驶和辅助驾驶技术的发展,对复杂驾驶场景的理解能力提出了更高要求。多模态通用大模型成为应对这一挑战的方案,但在垂直领域应用中仍面临数据采集、模型训练和部署优化等难题。本文提出一套面向驾驶场景的多模态模型优化方法,涵盖锥桶检测、交通灯识别、限速推荐和路口预警等任务。方法包括动态提示优化、数据集构建、模型训练与部署优化。动态提示根据输入图像内容自适应调整,强化模型对影响本车物体的关注力,提升任务专注度与判断能力。数据集通过融合真实与合成数据构建,实现高质量与多样性,增强模型在复杂环境中的泛化能力。模型训练阶段整合知识蒸馏、动态微调与量化技术,在降低存储与计算开销的同时提升性能。实验结果表明,该系统化优化方法显著提升了关键任务的准确率,并实现高效资源利用,为驾驶场景感知技术的实际应用提供有力支持。

原文摘要 · Abstract (English)

With the advancement of autonomous and assisted driving technologies, higher demands are placed on the ability to understand complex driving scenarios. Multimodal general large models have emerged as a solution for this challenge. However, applying these models in vertical domains involves difficulties such as data collection, model training, and deployment optimization. This paper proposes a comprehensive method for optimizing multimodal models in driving scenarios, including cone detection, traffic light recognition, speed limit recommendation, and intersection alerts. The method covers key aspects such as dynamic prompt optimization, dataset construction, model training, and deployment. Specifically, the dynamic prompt optimization adjusts the prompts based on the input image content to focus on objects affecting the ego vehicle, enhancing the model's task-specific focus and judgment capabilities. The dataset is constructed by combining real and synthetic data to create a high-quality and diverse multimodal training dataset, improving the model's generalization in complex driving environments. In model training, advanced techniques like knowledge distillation, dynamic fine-tuning, and quantization are integrated to reduce storage and computational costs while boosting performance. Experimental results show that this systematic optimization method not only significantly improves the model's accuracy in key tasks but also achieves efficient resource utilization, providing strong support for the practical application of driving scenario perception technologies.

自动驾驶多模态提示优化模型压缩

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。