arXiv:2608.28762cs.CV2026-08中稿 · EMNLP

构建3D路口多模态问答基准,提升交通场景时空推理能力

Inter-3D VQA: A Roadside Multimodal Benchmark for 3D Spatiotemporally Grounded Visual Question Answering

论文配图:Inter-3D VQA: A Roadside Multimodal Benchmark for 3D Spatiotemporally Grounded Visual Question Answering
图 1 · 摘自论文原文
  • 基于同步点云与多视角图像构建3D时空定位问答数据集
  • 包含40.7万条问答对,覆盖车道级位置与近距交互推理
  • 提出融合激光雷达表征的模型与统一评估框架,适合自动驾驶研究

视觉问答与多模态大语言模型的进展使得自然语言推理成为可能。然而,现有基准主要基于车载视角或2D路侧视频,难以评估真实距离、轨迹、基础设施拓扑及安全关键交互中的3D空间时序推理。本文提出Inter-3D VQA,一个大规模路侧多模态基准,用于交叉口的3D时空定位视觉问答。该数据集基于同步点云与多视角图像构建,包含407,000个问答对,涵盖车道级位置、物体关系、运动模式及近距交互推理。我们进一步提出Inter-Geo(融合物体与场景级对齐激光雷达表示的MLLM基线)和Inter-Metrics(统一评估文本一致性、数值准确性和语义正确性的框架)。实验表明,Inter-Geo在3D空间与时间推理任务上优于基于图像的VLM。代码与数据已开源。

原文摘要 · Abstract (English)

Recent advances in visual question answering (VQA) and multimodal large language models (MLLMs) have enabled natural-language reasoning over traffic scenes. However, existing benchmarks are largely built from ego-vehicle views or 2D roadside videos, limiting their ability to evaluate 3D-grounded reasoning over real-world distances, trajectories, infrastructure topology, and safety-critical interactions. We introduce Inter-3D VQA, a large-scale roadside multimodal benchmark for 3D spatiotemporally grounded VQA at intersections. Built from synchronized point clouds and multi-view images, Inter-3D VQA contains 407K QA pairs covering lane-level positions, object relationships, motion patterns, and near-miss-oriented interaction reasoning. We further propose Inter-Geo, an MLLM baseline that integrates object- and scene-level aligned LiDAR representations, and Inter-Metrics, a unified evaluation framework for textual consistency, numerical accuracy, and semantic correctness. Experiments show that Inter-Geo outperforms image-based VLMs, especially on grounded spatial and temporal reasoning tasks. Our benchmark and codes are available at https://github.com/ASU-Suo-Lab/Inter-3D-VQA .

3D视觉问答自动驾驶多模态激光雷达

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。