arXiv:2512.24605cs.CV2025-12被引 1

构建首个路侧设备采集的3D视觉定位数据集,支持交通场景语言理解。

MoniRefer: A Real-world Large-scale Multi-modal Dataset based on Roadside Infrastructure for 3D Visual Grounding

  • 基于路侧传感器采集真实交通路口点云与文本描述,构建大规模多模态数据集。
  • 包含约13.6万物体和41万条自然语言表达,经人工验证确保标注质量。
  • 提出端到端模型Moni3DVG,融合图像与点云信息提升定位精度,适合智能交通研究者。

3D视觉定位旨在从点云场景中定位与自然语言描述语义对应的目标物体,对路侧基础设施系统理解复杂交通环境至关重要。然而,现有3D视觉定位数据集和方法多聚焦于室内或车载视角下的驾驶场景,缺乏由路侧设备采集、面向户外监控任务的配对点云-文本数据。为此,本文提出面向户外监控场景的3D视觉定位新任务,实现超越车辆自身视角的基础设施级场景理解。我们构建了首个真实世界大规模多模态数据集MoniRefer,涵盖多个复杂交通路口的真实环境数据,共包含约136,018个物体及411,128条自然语言表达。所有语言描述与3D标注均经人工验证以确保质量。此外,我们提出一种新的端到端方法Moni3DVG,利用图像提供的丰富外观信息以及点云中的几何与光度信息进行多模态特征学习与3D目标定位。在所提基准上的大量实验与消融研究证明了该方法的有效性与优越性。数据集与代码将公开发布。

原文摘要 · Abstract (English)

3D visual grounding aims to localize the object in 3D point cloud scenes that semantically corresponds to given natural language sentences. It is very critical for roadside infrastructure system to interpret natural languages and localize relevant target objects in complex traffic environments. However, most existing datasets and approaches for 3D visual grounding focus on the indoor and outdoor driving scenes, outdoor monitoring scenarios remain unexplored due to scarcity of paired point cloud-text data captured by roadside infrastructure sensors. In this paper, we introduce a novel task of 3D Visual Grounding for Outdoor Monitoring Scenarios, which enables infrastructure-level understanding of traffic scenes beyond the ego-vehicle perspective. To support this task, we construct MoniRefer, the first real-world large-scale multi-modal dataset for roadside-level 3D visual grounding. The dataset consists of about 136,018 objects with 411,128 natural language expressions collected from multiple complex traffic intersections in the real-world environments. To ensure the quality and accuracy of the dataset, we manually verified all linguistic descriptions and 3D labels for objects. Additionally, we also propose a new end-to-end method, named Moni3DVG, which utilizes the rich appearance information provided by images and geometry and optical information from point cloud for multi-modal feature learning and 3D object localization. Extensive experiments and ablation studies on the proposed benchmarks demonstrate the superiority and effectiveness of our method. Our dataset and code will be released.

3D视觉定位多模态路侧感知交通监控

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。