BADAS-2.0提升碰撞预警能力,支持实时部署与可解释性。
Beyond the Beep: Scalable Collision Anticipation and Real-Time Explainability with BADAS-2.0
- 用自监督预训练+知识蒸馏,生成轻量级模型实现毫秒级推理。
- 构建17.8万段标注视频的长尾数据集,显著提升罕见高危场景识别率。
- 输出注意力热图与自然语言推理,让预测结果可理解、可信任。
我们提出BADAS-2.0,第二代碰撞预警系统,在前代基础上实现三方面突破:(i) 构建10类长尾场景基准,利用BADAS-1.0作为主动标注代理筛选百万段未标注驾驶视频,结合Nexar Atlas平台定向采集数据,将标注视频从4万增至17.85万(约200万片段),在所有子组中均取得提升,尤其在最难长尾类别上表现显著;(ii) 基于225万段无标签驾驶视频进行领域自监督预训练,通过知识蒸馏生成轻量模型BADAS-2.0-Flash(86M)和BADAS-2.0-Flash-Lite(22M),实现7–12倍速度提升且精度接近原模型,支持实时边缘部署;(iii) 生成实时对象中心注意力热图定位预测依据,进一步通过视觉-语言模型构建BADAS-Reason,基于最后一帧与热图生成驾驶员行为描述与结构化文本推理。代码与评估基准已公开。
原文摘要 · Abstract (English)
We present BADAS-2.0, the second generation of our collision anticipation system, building on BADAS-1.0, which showed that fine-tuning V-JEPA2 on large-scale ego-centric dashcam data outperforms both academic baselines and production ADAS systems. BADAS-2.0 advances the state of the art along three axes. (i) Long-tail benchmark and accuracy: We introduce a 10-group long-tail benchmark targeting rare and safety-critical scenarios. To construct it, BADAS-1.0 is used as an active oracle to score millions of unlabeled drives and surface high-risk candidates for annotation. Combined with Nexar's Atlas platform for targeted data collection, this expands the dataset from 40k to 178,500 labeled videos (~2M clips), yielding consistent gains across all subgroups, with the largest improvements on the hardest long-tail cases. (ii) Knowledge distillation to edge: Domain-specific self-supervised pre-training on 2.25M unlabeled driving videos enables distillation into compact models, BADAS-2.0-Flash (86M) and BADAS-2.0-Flash-Lite (22M), achieving 7-12x speedup with near-parity accuracy, enabling real-time edge deployment. (iii) Explainability: BADAS-2.0 produces real-time object-centric attention heatmaps that localize the evidence behind predictions. BADAS-Reason extends this with a vision-language model that consumes the last frame and heatmap to generate driver actions and structured textual reasoning. Inference code and evaluation benchmarks are publicly available.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。