arXiv:2606.01939cs.CV2026-06

用全景视频生成大型仓库的语义线框地图,精度达4.8厘米。

SAVMap: Structure-Aided Visual Mapping of Large-Scale 2.5D Manhattan Wireframes from Panoramic Video

论文配图:SAVMap: Structure-Aided Visual Mapping of Large-Scale 2.5D Manhattan Wireframes from Panoramic Video
图 1 · 摘自论文原文
  • 通过全景视频提取货架与天花板视角图像,结合语义分割定位关键结构点。
  • 在46排货架、55米×7米每排的仓库中,生成超5000个货架元素的线框图,误差仅4.8厘米。
  • 适合工业级机器人定位与数字孪生应用,无需额外传感器。

精确的三维环境表示可支持机器人定位与数字孪生等任务。本文提出SAVMap,仅使用全景视频摄像头作为输入,生成仓库货架与照明结构的语义线框地图。沿着仓库过道拍摄的全景视频被处理为一系列校正后的图像,涵盖货架与天花板朝向视图。通过前端语义分割网络,从每张图像中提取稀疏的语义结构特征点(如货架角点、灯具中心),并追踪其在序列中的变化。利用真实世界中点间存在的曼哈顿网格几何关系,采用约束式结构光流算法恢复出三维点云,构建线框地图。我们在一个包含46排货架、每排尺寸55米×7米的仓库中验证了该方法的可扩展性与精度:仅用一小时的全景视频内容,便成功构建了超过5000个货架元素的线框地图,整体平均绝对误差为4.8厘米,优于地面真值。

原文摘要 · Abstract (English)

Precise 3D representations of industrial environments enable tasks such as robot localization and digital twin generation. We propose SAVMap, a method for generating a semantic wireframe map of warehouse shelf and light structures using only a panoramic video camera as the sensor input. Sequences of rectified images with shelf and ceiling-facing views are extracted from a panoramic video captured along the warehouse aisles. Using a semantic segmentation network front end, a set of sparse, semantic structure feature points (e.g., corners of shelf structures, centers of lights) are extracted from each image and tracked across the sequences. By accounting for real-world geometric relationships among the points such as Manhattan grids, a constrained structure-from-motion algorithm yields the 3D points that form a wireframe map. We demonstrate the scalability and accuracy of our proposal in a warehouse with 46 shelving rows, each with faces spanning 55\,m by 7\,m. From an hour of panoramic video content, we create wireframe maps for over 5000 shelf elements across the rows, achieving an aggregate mean absolute error of 4.8\,cm with respect to ground-truth.

3D重建线框地图全景视频工业感知

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。