让视觉语言模型提前退出,提速近60%且更准
AD-EE: Early Exiting for Fast and Reliable Vision-Language Models in Autonomous Driving
- 根据自动驾驶场景特点,用因果推理选最优退出层
- 实测延迟最高降低57.58%,物体检测准确率最高提升44%
- 适合对实时性要求高的自动驾驶感知系统
随着自动驾驶技术快速发展,视觉语言模型(VLMs)被广泛用于提升感知与决策能力。然而,高延迟和计算开销严重制约其在时间敏感驾驶场景中的应用,尤其在模型过推理——即使已获得确定判断仍继续处理多余层级时更为明显。为此,我们提出AD-EE早期退出框架,融合自动驾驶领域特性并利用因果推断识别最优退出层。在大规模真实驾驶数据集Waymo和侧重极端场景的CODA上,以及搭载Autoware Universe平台的真实车辆上进行了评估。多模型实验表明,该方法显著降低延迟,最大改善达57.58%;同时提升目标检测精度,最大增益达44%。
原文摘要 · Abstract (English)
With the rapid advancement of autonomous driving, deploying Vision-Language Models (VLMs) to enhance perception and decision-making has become increasingly common. However, the real-time application of VLMs is hindered by high latency and computational overhead, limiting their effectiveness in time-critical driving scenarios. This challenge is particularly evident when VLMs exhibit over-inference, continuing to process unnecessary layers even after confident predictions have been reached. To address this inefficiency, we propose AD-EE, an Early Exit framework that incorporates domain characteristics of autonomous driving and leverages causal inference to identify optimal exit layers. We evaluate our method on large-scale real-world autonomous driving datasets, including Waymo and the corner-case-focused CODA, as well as on a real vehicle running the Autoware Universe platform. Extensive experiments across multiple VLMs show that our method significantly reduces latency, with maximum improvements reaching up to 57.58%, and enhances object detection accuracy, with maximum gains of up to 44%.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。