arXiv:2604.08008cs.CVcs.AI2026-04

构建大规模罕见驾驶场景图像检索数据集,助力自动驾驶系统感知长尾风险。

SearchAD: Large-Scale Rare Image Retrieval Dataset for Autonomous Driving

  • 从11个数据集整合42.3万帧图像,标注51.3万个边界框覆盖90类罕见场景。
  • 提出文本与图像双向检索任务,验证文本方法在语义理解上优于图像方法。
  • 专为长尾感知与数据筛选设计,适合研究少样本学习与多模态模型优化者。

从大规模数据集中检索罕见且关乎安全的驾驶场景,对构建鲁棒的自动驾驶系统至关重要。随着数据规模持续增长,核心挑战已从数据收集转向高效识别关键样本。我们提出SearchAD,一个面向自动驾驶的大规模罕见图像检索数据集,包含超过42.3万帧图像,源自11个现有数据集。该数据集提供超过51.3万个高质量人工标注的边界框,涵盖90种罕见类别,部分类别在整个数据集中出现次数少于50次。不同于以往聚焦实例级检索的基准,SearchAD强调语义级图像检索,并设有明确的数据划分,支持文本到图像、图像到图像检索、少样本学习及多模态检索模型微调。全面评估显示,基于文本的方法因更强的语义表征能力,性能优于基于图像的方法。尽管直接对齐空间视觉特征与语言的模型在零样本场景下表现最佳,且微调基线显著提升性能,但整体检索能力仍不理想。通过在公开基准服务器上的预留测试集,SearchAD成为首个面向检索驱动的数据清洗与长尾感知研究的大规模数据集:https://iis-esslingen.github.io/searchad/

原文摘要 · Abstract (English)

Retrieving rare and safety-critical driving scenarios from large-scale datasets is essential for building robust autonomous driving (AD) systems. As dataset sizes continue to grow, the key challenge shifts from collecting more data to efficiently identifying the most relevant samples. We introduce SearchAD, a large-scale rare image retrieval dataset for AD containing over 423k frames drawn from 11 established datasets. SearchAD provides high-quality manual annotations of more than 513k bounding boxes covering 90 rare categories. It specifically targets the needle-in-a-haystack problem of locating extremely rare classes, with some appearing fewer than 50 times across the entire dataset. Unlike existing benchmarks, which focused on instance-level retrieval, SearchAD emphasizes semantic image retrieval with a well-defined data split, enabling text-to-image and image-to-image retrieval, few-shot learning, and fine-tuning of multi-modal retrieval models. Comprehensive evaluations show that text-based methods outperform image-based ones due to stronger inherent semantic grounding. While models directly aligning spatial visual features with language achieve the best zero-shot results, and our fine-tuning baseline significantly improves performance, absolute retrieval capabilities remain unsatisfactory. With a held-out test set on a public benchmark server, SearchAD establishes the first large-scale dataset for retrieval-driven data curation and long-tail perception research in AD: https://iis-esslingen.github.io/searchad/

自动驾驶图像检索长尾问题多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。