用自然语言查询,精准找出行人复杂行为的驾驶场景。
Context-based Motion Retrieval using Open Vocabulary Methods for Autonomous Driving
- 结合人体模型与视频帧,用文本检索人类行为
- 在新数据集上比现有方法高27.5%准确率
- 适合评估自动驾驶对行人异常行为的应对能力
自动驾驶系统需在涉及脆弱道路使用者(VRUs)的罕见或复杂行为场景中可靠运行。识别这些边缘案例对系统评估与泛化至关重要,但从大规模数据集中检索这类稀有行为仍具挑战。为此,我们提出一种上下文感知的运动检索框架,将基于SMPL的人体运动序列与对应视频帧融合编码至共享多模态嵌入空间,并与自然语言对齐,实现通过文本查询高效检索人类行为及其上下文。本工作还引入了新数据集WayMoCo,作为Waymo Open Dataset的扩展,包含由生成的伪真值SMPL序列和对应图像数据自动生成的运动与场景描述标注。在WayMoCo数据集上,该方法相较当前最优模型在运动-上下文检索任务中最高提升27.5%的准确率。
原文摘要 · Abstract (English)
Autonomous driving systems must operate reliably in safety-critical scenarios, particularly those involving unusual or complex behavior by Vulnerable Road Users (VRUs). Identifying these edge cases in driving datasets is essential for robust evaluation and generalization, but retrieving such rare human behavior scenarios within the long tail of large-scale datasets is challenging. To support targeted evaluation of autonomous driving systems in diverse, human-centered scenarios, we propose a novel context-aware motion retrieval framework. Our method combines Skinned Multi-Person Linear (SMPL)-based motion sequences and corresponding video frames before encoding them into a shared multimodal embedding space aligned with natural language. Our approach enables the scalable retrieval of human behavior and their context through text queries. This work also introduces our dataset WayMoCo, an extension of the Waymo Open Dataset. It contains automatically labeled motion and scene context descriptions derived from generated pseudo-ground-truth SMPL sequences and corresponding image data. Our approach outperforms state-of-the-art models by up to 27.5% accuracy in motion-context retrieval, when evaluated on the WayMoCo dataset.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。