用单张参考图实现车载实时异常检测,无需标注数据。
Real-World On-Vehicle Evaluation of Embedding-Based Anomaly Detection

- 基于预训练视觉变换器嵌入,通过特征空间近邻相似性检测异常
- 仅需一张正常图像即可建模正常状态,在真实道路表现稳定
- 适合无监督部署,特别适用于自动驾驶真实场景
交通场景中的异常检测对自动驾驶安全至关重要,但获取代表性异常数据仍具挑战。现有方法高度依赖城市景观语义类定义的正常性,难以适应多样化真实场景。本文提出一种可适配的实时异常检测方法,利用预训练视觉变换器嵌入,在潜在语义特征空间中通过最近邻相似性检测偏差。基于局部块处理,算法生成密集异常掩码,实现异常定位。该方法仅需单张参考图像即可稳健建模正常性,避免显式监督与数据集特定训练,适合真实世界部署。我们在标准基准和真实车辆上评估了该方法。尽管结构简单,其在Road Anomaly基准上表现良好,并在实际场景中展现出一致的定性行为,成功识别出多种场景下的语义异常物体。结果表明,简单的参考式方法在真实运行条件下仍能提供有效异常信号。
原文摘要 · Abstract (English)
Detecting anomalies in traffic scenes is crucial for ensuring safety in autonomous driving, yet collecting representative anomalous data remains challenging. Existing anomaly detection methods are highly specialized and rely on normality as defined by the abstract semantic Cityscapes classes, making it difficult to adapt to diverse real-world scenarios. We propose an adaptable real-time anomaly detection method that leverages foundation models in the form of pretrained vision transformer embeddings to detect deviations via nearest-neighbor similarity in the latent semantic feature space. Based on patch-wise processing, the algorithm produces dense anomaly masks, allowing for the localization of detected anomalies. The method robustly models normality through a single reference image. This formulation avoids explicit supervision and dataset-specific training, making it suitable for real-world deployment. We evaluate the method on standard benchmarks and on an automated vehicle in real-world scenarios. Despite its simplicity, the method achieves good performance on the Road Anomaly benchmark and demonstrates consistent qualitative behavior in practice, successfully highlighting semantically unusual objects in diverse scenes. These results suggest that simple, reference-based methods can provide useful anomaly signals under realistic operating conditions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。