arXiv:2411.12620cs.CV2024-11被引 1

用稀疏多视角图像自动生成2D语义地图,精度优于GPS。

Maps from Motion (MfM): Generating 2D Semantic Maps from Sparse Multi-view Images

  • 基于图结构融合多视角检测结果,构建全局语义地图。
  • 在视点变化大、数据稀疏时平均定位误差仅4米,低于GPS精度。
  • 适合自动驾驶与地图自动化更新,尤其适用于难配准场景。

全球详尽的2D地图需耗费巨大人力。以OpenStreetMap为例,1100万用户手动标注超过17.5亿条地理信息,包括标志性建筑和常见城市物体。但人工标注易出错且更新缓慢,影响地图准确性。本文提出Maps from Motion(MfM),通过一组未标定的多视角图像直接生成2D语义地图,从每张图像中提取目标检测,并估计其在相机参考系下的俯视局部地图。由于局部地图不完整、噪声大,且城市物体外观相似导致匹配不可靠,对齐过程极具挑战。为此,我们提出一种新型图基框架,编码各图像中目标的空间与语义分布,学习如何组合以预测全局坐标系下的目标姿态,同时考虑所有可能的检测匹配并保留图像内的拓扑结构。尽管问题复杂,最优模型在稀疏序列且强视角变化下仍实现平均4米以内全局注册精度,远超标准方法;在COLMAP失败率达80%的场景中仍有效。我们在合成与真实数据上进行了广泛评估,证明该方法在传统优化失败的场景下仍能获得可行解。

原文摘要 · Abstract (English)

World-wide detailed 2D maps require enormous collective efforts. OpenStreetMap is the result of 11 million registered users manually annotating the GPS location of over 1.75 billion entries, including distinctive landmarks and common urban objects. At the same time, manual annotations can include errors and are slow to update, limiting the map's accuracy. Maps from Motion (MfM) is a step forward to automatize such time-consuming map making procedure by computing 2D maps of semantic objects directly from a collection of uncalibrated multi-view images. From each image, we extract a set of object detections, and estimate their spatial arrangement in a top-down local map centered in the reference frame of the camera that captured the image. Aligning these local maps is not a trivial problem, since they provide incomplete, noisy fragments of the scene, and matching detections across them is unreliable because of the presence of repeated pattern and the limited appearance variability of urban objects. We address this with a novel graph-based framework, that encodes the spatial and semantic distribution of the objects detected in each image, and learns how to combine them to predict the objects' poses in a global reference system, while taking into account all possible detection matches and preserving the topology observed in each image. Despite the complexity of the problem, our best model achieves global 2D registration with an average accuracy within 4 meters (i.e., below GPS accuracy) even on sparse sequences with strong viewpoint change, on which COLMAP has an 80% failure rate. We provide extensive evaluation on synthetic and real-world data, showing how the method obtains a solution even in scenarios where standard optimization techniques fail.

语义地图多视角重建自动驾驶图神经网络

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。