arXiv:2505.15379cs.CV2025-05

P³数据集融合点云与影像,提升建筑矢量化精度。

The P$^3$ Dataset: Pixels, Points and Polygons for Multimodal Building Vectorization

  • 结合航拍点云、影像与矢量轮廓,构建多模态数据集。
  • 点云在混合与端到端框架中均提升建筑多边形预测效果。
  • 融合点云与影像可显著改善生成多边形的几何质量。

我们提出P³数据集,一个大规模多模态建筑矢量化基准,涵盖三个大洲的航空激光雷达点云、高分辨率航拍影像及二维建筑矢量轮廓。数据集包含超过100亿个激光雷达点,精度达分米级,影像地面采样距离为25厘米。现有数据集多聚焦于图像模态,而P³补充了密集3D信息。实验表明,激光雷达点云在混合与端到端学习框架中均能有效预测建筑多边形;进一步融合点云与影像可显著提升预测多边形的准确性和几何质量。该数据集已公开,配套代码与三个先进模型的预训练权重可在https://github.com/raphaelsulzer/PixelsPointsPolygons 获取。

原文摘要 · Abstract (English)

We present the P$^3$ dataset, a large-scale multimodal benchmark for building vectorization, constructed from aerial LiDAR point clouds, high-resolution aerial imagery, and vectorized 2D building outlines, collected across three continents. The dataset contains over 10 billion LiDAR points with decimeter-level accuracy and RGB images at a ground sampling distance of 25 centimeter. While many existing datasets primarily focus on the image modality, P$^3$ offers a complementary perspective by also incorporating dense 3D information. We demonstrate that LiDAR point clouds serve as a robust modality for predicting building polygons, both in hybrid and end-to-end learning frameworks. Moreover, fusing aerial LiDAR and imagery further improves accuracy and geometric quality of predicted polygons. The P$^3$ dataset is publicly available, along with code and pretrained weights of three state-of-the-art models for building polygon prediction at https://github.com/raphaelsulzer/PixelsPointsPolygons .

建筑矢量化多模态激光雷达点云

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。