构建首个面向印度道路场景的多任务多模态问答数据集
RoadscapesQA: A Multitask, Multimodal Dataset for Visual Question Answering on Indian Roads
- 用规则启发式自动生成带标注的视觉问答对
- 涵盖9000张图像,覆盖城乡、昼夜多种复杂路况
- 适合研究自动驾驶中非结构化环境理解的学者
理解道路场景对自动驾驶至关重要,可帮助系统解读视觉环境以支持有效决策。我们提出Roadscapes,一个包含最多9000张在多样化印度驾驶环境中拍摄的图像的多任务多模态数据集,并配有手动验证的边界框。为促进可扩展的场景理解,我们采用基于规则的启发式方法推断多种场景属性,并据此生成用于目标定位、推理和场景理解等任务的问答对。数据集涵盖城市与乡村道路,包括高速公路、服务区道路、村庄小路及拥堵城区街道,拍摄于白天与夜间。Roadscapes旨在推动非结构化环境中视觉场景理解的研究。本文描述了数据收集与标注流程,呈现关键数据集统计,并提供基于视觉-语言模型的图像问答任务初步基线。
原文摘要 · Abstract (English)
Understanding road scenes is essential for autonomous driving, as it enables systems to interpret visual surroundings to aid in effective decision-making. We present Roadscapes, a multitask multimodal dataset consisting of upto 9,000 images captured in diverse Indian driving environments, accompanied by manually verified bounding boxes. To facilitate scalable scene understanding, we employ rule-based heuristics to infer various scene attributes, which are subsequently used to generate question-answer (QA) pairs for tasks such as object grounding, reasoning, and scene understanding. The dataset includes a variety of scenes from urban and rural India, encompassing highways, service roads, village paths, and congested city streets, captured in both daytime and nighttime settings. Roadscapes has been curated to advance research on visual scene understanding in unstructured environments. In this paper, we describe the data collection and annotation process, present key dataset statistics, and provide initial baselines for image QA tasks using vision-language models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。