arXiv:2606.28383cs.CVcs.LG2026-06被引 1

无需标签,用预测误差自动识别复杂驾驶场景。

Zero-Label Driving Scenario Complexity Detection via Joint Embedding Predictive Architecture

论文配图:Zero-Label Driving Scenario Complexity Detection via Joint Embedding Predictive Architecture
图 1 · 摘自论文原文
  • 用自监督模型预测车辆状态,通过误差判断场景复杂度。
  • 在无标签数据中准确识别出行人交互等高危场景。
  • 适合自动驾驶系统中的安全评估与异常检测应用。

在大规模未标注数据集中识别复杂且具有安全风险的驾驶场景是一个重要但成本高昂的问题。现有方法依赖人工标注、有监督分类器或精心设计的规则集,均需对困难场景有先验知识。我们探讨模型能否在无任何标签的情况下自主发现场景复杂度。在nuPlan mini数据集的结构化代理状态数据上训练一个极简的联合嵌入预测架构(JEPA),并使用时间预测误差作为零样本复杂度评分。训练和评估过程中均未访问真实标签,模型对涉及无保护转弯、人行横道互动和行人接近的场景赋予显著更高的评分,而对车道保持和静止交通场景则赋予显著更低的评分。通过四项消融实验隔离信号来源,并通过下游异常检测评估验证:平均精度达0.512,优于0.436的随机基线。结果表明,自监督潜在世界模型中的时间预测误差可作为驾驶场景复杂度的有效代理。

原文摘要 · Abstract (English)

Identifying complex and safety-critical driving scenarios in large unlabelled datasets is an important but expensive problem. Existing approaches rely on human annotators, supervised classifiers, or carefully engineered rule sets, all of which require substantial prior knowledge about what constitutes a difficult scenario. We ask whether a model can discover scenario complexity on its own, with no labels at any stage. We train a minimal Joint Embedding Predictive Architecture (JEPA) on structured agent state data from the nuPlan mini dataset and use the temporal prediction error as a zero-shot complexity score. Without access to any ground-truth labels during training or evaluation setup, the model assigns significantly higher scores to scenarios involving unprotected turns, crosswalk interactions, and pedestrian proximity, and significantly lower scores to lane-following and stationary-traffic scenarios. We validate this finding through four ablation experiments that isolate the source of the signal, and through a downstream anomaly detection evaluation that achieves Average Precision of 0.512 against a 0.436 chance baseline. The results show that temporal prediction error in a self-supervised latent world model is a practical proxy for driving scenario complexity.

自动驾驶自监督学习场景复杂度零样本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。