arXiv:2505.20981cs.CVcs.CL2025-05被引 17

用自然语言精准定位自动驾驶中的复杂交通场景。

RefAV: Towards Planning-Centric Scenario Mining

  • 基于视觉语言模型,将自然语言查询映射到驾驶日志的时空位置。
  • 构建包含1万条查询的RefAV数据集,覆盖Argoverse 2中1000段真实驾驶日志。
  • 发现通用VLM直接迁移效果差,凸显场景挖掘的独特挑战。

自动驾驶车辆在常规车队测试中收集并伪标注了大量多模态数据,这些数据与高精地图关联。然而,从未经整理的驾驶日志中识别出有趣且安全关键的场景仍具挑战性。传统场景挖掘方法易出错且耗时,常依赖人工设计的结构化查询。本文通过近期视觉语言模型(VLMs)的视角重新审视时空场景挖掘,旨在判断描述的场景是否出现在驾驶日志中,并精确定位其时空位置。为此,我们构建了RefAV——一个大规模数据集,包含10,000条源自Argoverse 2 Sensor数据集中1000段驾驶日志的自然语言查询,涵盖与运动规划相关的复杂多智能体交互。我们评估了多种参照式多对象追踪器,并对基线进行实证分析。值得注意的是,简单复用现成VLM性能不佳,表明场景挖掘具有独特挑战。最后,我们介绍了近期举办的竞赛及社区反馈。代码与数据集详见:https://github.com/CainanD/RefAV/ 与 https://argoverse.github.io/user-guide/tasks/scenario_mining.html。

原文摘要 · Abstract (English)

Autonomous Vehicles (AVs) collect and pseudo-label terabytes of multi-modal data localized to HD maps during normal fleet testing. However, identifying interesting and safety-critical scenarios from uncurated driving logs remains a significant challenge. Traditional scenario mining techniques are error-prone and prohibitively time-consuming, often relying on hand-crafted structured queries. In this work, we revisit spatio-temporal scenario mining through the lens of recent vision-language models (VLMs) to detect whether a described scenario occurs in a driving log and, if so, precisely localize it in both time and space. To address this problem, we introduce RefAV, a large-scale dataset of 10,000 diverse natural language queries that describe complex multi-agent interactions relevant to motion planning derived from 1000 driving logs in the Argoverse 2 Sensor dataset. We evaluate several referential multi-object trackers and present an empirical analysis of our baselines. Notably, we find that naively repurposing off-the-shelf VLMs yields poor performance, suggesting that scenario mining presents unique challenges. Lastly, we discuss our recently held competition and share insights from the community. Our code and dataset are available at https://github.com/CainanD/RefAV/ and https://argoverse.github.io/user-guide/tasks/scenario_mining.html

自动驾驶场景挖掘视觉语言模型多智能体

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。