构建道路场景VQA数据集,让智能交通系统能理解交通行为背后的意图与规则。
RoadSceneVQA: Benchmarking Visual Question Answering in Roadside Perception Systems for Intelligent Transportation System
- 设计道路场景专用的视觉问答数据集,涵盖3万+真实条件下的问答对。
- 提出CogniAnchor融合与AD-CoT推理框架,显著提升模型的逻辑推理能力。
- 适合关注智能交通、多模态推理与人车交互的科研与工程人员。
当前路边感知系统多聚焦于实例级识别,难以通过自然语言交互或上下文推理理解交通行为。为弥补这一缺口,我们提出RoadSceneVQA,一个大规模、精细化标注的视觉问答数据集,专用于路边场景。该数据集包含34,736组多样化的问答对,覆盖不同天气、光照和交通条件,不仅关注物体属性,还涉及交通参与者的行为意图、合法性及互动模式。任务要求模型结合现实交通规则与上下文依赖,完成显性识别与隐性常识推理。为充分发挥多模态大模型(MLLMs)潜力,我们提出受人类场景锚定机制启发的CogniAnchor Fusion(CAF)视觉-语言融合模块,并设计辅助解耦思维链(AD-CoT)以增强推理过程。基于上述方法,我们构建基准模型RoadMind。在RoadSceneVQA与CODA-LM基准测试中,该方案持续提升推理准确率与计算效率,使MLLM在结构化交通感知与推理任务中达到当前最优性能。
原文摘要 · Abstract (English)
Current roadside perception systems mainly focus on instance-level perception, which fall short in enabling interaction via natural language and reasoning about traffic behaviors in context. To bridge this gap, we introduce RoadSceneVQA, a large-scale and richly annotated visual question answering (VQA) dataset specifically tailored for roadside scenarios. The dataset comprises 34,736 diverse QA pairs collected under varying weather, illumination, and traffic conditions, targeting not only object attributes but also the intent, legality, and interaction patterns of traffic participants. RoadSceneVQA challenges models to perform both explicit recognition and implicit commonsense reasoning, grounded in real-world traffic rules and contextual dependencies. To fully exploit the reasoning potential of Multi-modal Large Language Models (MLLMs), we further propose CogniAnchor Fusion (CAF), a vision-language fusion module inspired by human-like scene anchoring mechanisms. Moreover, we propose the Assisted Decoupled Chain-of-Thought (AD-CoT) to enhance the reasoned thinking via CoT prompting and multi-task learning. Based on the above, we propose the baseline model RoadMind. Experiments on RoadSceneVQA and CODA-LM benchmark show that the pipeline consistently improves both reasoning accuracy and computational efficiency, allowing the MLLM to achieve state-of-the-art performance in structural traffic perception and reasoning tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。