融合视觉与轨迹信息,提升自动驾驶场景检索准确率。
Multimodal Scenario Similarity Search for Autonomous Driving

- 构建统一检索框架,联合使用视觉与轨迹特征。
- 轨迹特征在变道、转弯等动作场景中表现更优。
- 多模态融合效果最佳,适合数据挖掘与验证场景。
大规模自动驾驶数据集包含海量记录场景,亟需高效检索方法以定位与查询相似的驾驶情境。现有方法多依赖视觉表征或运动描述,难以全面评估其优劣。本文提出一种多模态自动驾驶场景检索框架,将视觉与基于轨迹的表示统一整合至同一检索流程中。研究对比两种轨迹方法:基于周围车辆运动显式匹配的Exo-Trajectory,以及通过对比学习从物体轨迹中学习的ScenarioFormer(Transformer架构)。在多种驾驶场景下与强视觉基线比较,结果表明:轨迹表征在变道、转弯及车流排队等以运动为核心的事件中表现优异;而视觉嵌入在外观线索丰富时更具优势。最重要的是,融合视觉与轨迹信息可持续提升检索质量,取得最优整体性能。实验验证了外观与运动是互补的场景相似性维度,推动了多模态检索系统在自动驾驶数据挖掘、数据集整理与场景化验证中的应用。
原文摘要 · Abstract (English)
Large-scale autonomous-driving datasets contain vast numbers of recorded scenarios, creating a need for efficient retrieval methods that can identify situations similar to a given query. Existing approaches typically rely on either visual representations or motion-based descriptions, making it difficult to understand their relative strengths and limitations for scenario retrieval. In this work, we present a multimodal framework for autonomous-driving scenario retrieval that combines visual and trajectory-based representations within a unified retrieval pipeline. We investigate two trajectory-based approaches: Exo-Trajectory, an explicit matching method based on surrounding-agent motion, and ScenarioFormer, a transformer-based representation learned from object trajectories using contrastive learning. We compare these approaches against strong vision-based baselines and analyze their behavior across a diverse set of driving scenarios. Experimental results show that trajectory representations provide strong retrieval performance for motion-centric events such as cut-ins, turning maneuvers, and traffic queueing, while visual embeddings excel when appearance cues are informative. Most importantly, combining visual and trajectory information consistently improves retrieval quality, yielding the best overall performance. These findings demonstrate that appearance and motion capture are complementary notions of scenario similarity and motivate multimodal retrieval systems for autonomous-driving data mining, dataset curation, and scenario-based validation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。