arXiv:2501.05081cs.LG2025-01被引 1

用小模型让视觉语言模型在自动驾驶中落地

DriVLM: Domain Adaptation of Vision-Language Models in Autonomous Driving

  • 用小型多模态模型适配自动驾驶场景
  • 在资源受限下仍保持有效推理能力
  • 适合边缘设备部署的智能驾驶研究者

近年来,大语言模型表现出色,推动了人工智能的发展,其参数规模和性能持续增长。特别是多模态大语言模型(MLLM)能够融合图像、视频、声音、文本等多种模态,在各类任务中展现出巨大潜力。然而,大多数MLLM需要极高计算资源,对多数研究者和开发者构成挑战。本文探索了小规模MLLM的实用性,并将其应用于自动驾驶领域,旨在推动MLLM在真实场景中的应用。

原文摘要 · Abstract (English)

In recent years, large language models have had a very impressive performance, which largely contributed to the development and application of artificial intelligence, and the parameters and performance of the models are still growing rapidly. In particular, multimodal large language models (MLLM) can combine multiple modalities such as pictures, videos, sounds, texts, etc., and have great potential in various tasks. However, most MLLMs require very high computational resources, which is a major challenge for most researchers and developers. In this paper, we explored the utility of small-scale MLLMs and applied small-scale MLLMs to the field of autonomous driving. We hope that this will advance the application of MLLMs in real-world scenarios.

自动驾驶多模态模型小模型领域适配

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。