让自动驾驶模型提前感知碰撞,提升安全性和驾驶得分。
Collision-Aware Vision-Language Learning for End-to-End Driving with Multimodal Infraction Datasets
- 用视频语言联合学习捕捉时序碰撞信号,实现主动预测。
- 在模拟和真实数据上均显著提升碰撞检测效果,驱动分提高14.12%。
- 适用于需高安全性的端到端自动驾驶系统,尤其关注碰撞防护。
高违规率仍是端到端自动驾驶的主要瓶颈,如CARLA Leaderboard上的低驾驶得分所示。尽管碰撞相关违规是闭环评估中的主要失败模式,但针对碰撞的表征学习仍受关注不足。为此,我们首先构建了视频-语言增强异常检测器(VLAAD),采用多实例学习(MIL)框架,获得稳定且时序定位准确的碰撞信号,实现主动预测。为将此能力引入闭环仿真,必须克服现有模拟器数据集缺乏多模态信息、场景单一的问题。因此,我们提出CARLA-Collide,一个大规模多模态数据集,覆盖多样道路网络中的真实碰撞事件。基于该数据训练的VLAAD可作为插件模块,无缝集成至现有端到端驾驶模型。将其接入预训练TransFuser++代理后,驾驶得分相对提升14.12%,仅需少量微调。此外,我们在开环设置下评估了VLAAD在真实世界数据上的泛化能力,为此构建了真实场景的多模态数据集Real-Collide,包含多样化行车记录仪视频与语义丰富标注。即使仅有0.6B参数,VLAAD在该基准上仍超越数十亿参数的视觉语言模型,AUC提升23.3%。
原文摘要 · Abstract (English)
High infraction rates remain the primary bottleneck for end-to-end (E2E) autonomous driving, as evidenced by the low driving scores on the CARLA Leaderboard. Despite collision-related infractions being the dominant failure mode in closed-loop evaluations, collision-aware representation learning has received limited attention. To address this gap, we first develop a Video-Language-Augmented Anomaly Detector (VLAAD), leveraging a Multiple Instance Learning (MIL) formulation to obtain stable, temporally localized collision signals for proactive prediction. To transition these capabilities into closed-loop simulations, we must overcome the limitations of existing simulator datasets, which lack multimodality and are frequently restricted to simple intersection scenarios. Therefore, we introduce CARLA-Collide, a large-scale multimodal dataset capturing realistic collision events across highly diverse road networks. Trained on this diverse simulator data, VLAAD serves as a collision-aware plug-in module that can be seamlessly integrated into existing E2E driving models. By integrating our module into a pretrained TransFuser++ agent, we demonstrate a 14.12% relative increase in driving score with minimal fine-tuning. Beyond closed-loop evaluation, we further assess the generalization capability of VLAAD in an open-loop setting using real-world driving data. To support this analysis, we introduce Real-Collide, a multimodal dataset of diverse dashcam videos paired with semantically rich annotations for collision detection and prediction. On this benchmark, despite containing only 0.6B parameters, VLAAD outperforms a multi-billion-parameter vision-language model, achieving a 23.3% improvement in AUC.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。