arXiv:2606.09109cs.CVcs.IR2026-06被引 1

用数据校准规则,让自动驾驶视频检索更准地找到复杂事件。

Driving Video Retrieval for Complex Queries with Structured Grounding

论文配图:Driving Video Retrieval for Complex Queries with Structured Grounding
图 1 · 摘自论文原文
  • 通过弱标注数据评估规则可靠性并动态调整
  • 在三个基准上提升84%的首条准确率
  • 适合需要精准检索驾驶事件的研究者

大规模视频检索对自动驾驶的数据整理与安全验证至关重要,用户不仅需找场景,还需定位变道、急刹车等动态事件。现有视觉语言和关键词检索方法常遗漏这些事件,因相关运动未被文本明确描述或缺乏词汇重叠。基于规则的方法虽可直接编码事件,但易受现实数据偏差影响而失效。我们提出STRIVE-D,一种数据校准的驾驶视频检索框架:利用领域内弱标注视频估计查询规则的可靠性,自适应调整不匹配规则,并融合校准后的规则得分与视觉语言、关键词检索信号。在三个驾驶基准(包括新发布的DrivingDojo人类标注事件数据集)上,STRIVE-D相比最先进方法在Top-1准确率上最高提升84%。

原文摘要 · Abstract (English)

Video retrieval at scale is central to data curation and safety validation in autonomous driving, where users want to find not only scenes but also dynamic events such as cut-ins and hard braking. Existing vision-language and keyword-based retrieval methods often miss these events because the relevant motion may not be explicitly described in text or captured by lexical overlap. Rule-based retrieval can encode such events more directly, but it is brittle: generated or hand-written rules often fail when their assumptions do not match real driving data. We propose STRIVE-D, a data-calibrated retrieval framework for driving videos. It uses weakly labeled in-domain videos to estimate when a query rule is reliable, adapt rules that mismatch observed data, and fuse calibrated rule scores with vision-language and keyword-based retrieval signals. Across three driving benchmarks, including newly released human-annotated event data on DrivingDojo, STRIVE-D delivers up to 84% relative improvement in top-1 accuracy over state-of-the-art methods.

视频检索自动驾驶规则校准事件识别

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。