构建真实场景下车辆与人交互的指令数据集,支持自然语言导航
doScenes: An Autonomous Driving Dataset with Natural Language Instruction for Human Interaction and Vision-Language Navigation
- 基于真实驾驶数据标注自然语言指令与物体参照关系
- 支持对静态和动态物体的即时动作指令响应
- 适合研究人车协同、视觉语言导航的算法开发者
自动驾驶系统需将人类指令有效融入运动规划。本文提出doScenes数据集,聚焦短期直接影响车辆运动的指令,通过为多模态传感器数据标注自然语言指令及指代性标签,建立指令与驾驶行为间的桥梁,实现上下文感知与自适应规划。不同于仅关注排序或场景级推理的现有数据集,doScenes强调与静态和动态物体相关的可执行指令。该框架克服了以往研究依赖模拟数据或预定义动作集的局限,支持在真实场景中进行细致灵活的响应。本工作为开发无缝融合人类指令的自主系统学习策略奠定基础,推动视觉-语言导航中安全高效的人车协作。数据已开源:https://www.github.com/rossgreer/doScenes
原文摘要 · Abstract (English)
Human-interactive robotic systems, particularly autonomous vehicles (AVs), must effectively integrate human instructions into their motion planning. This paper introduces doScenes, a novel dataset designed to facilitate research on human-vehicle instruction interactions, focusing on short-term directives that directly influence vehicle motion. By annotating multimodal sensor data with natural language instructions and referentiality tags, doScenes bridges the gap between instruction and driving response, enabling context-aware and adaptive planning. Unlike existing datasets that focus on ranking or scene-level reasoning, doScenes emphasizes actionable directives tied to static and dynamic scene objects. This framework addresses limitations in prior research, such as reliance on simulated data or predefined action sets, by supporting nuanced and flexible responses in real-world scenarios. This work lays the foundation for developing learning strategies that seamlessly integrate human instructions into autonomous systems, advancing safe and effective human-vehicle collaboration for vision-language navigation. We make our data publicly available at https://www.github.com/rossgreer/doScenes
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。