arXiv:2506.15718cs.LG2025-06

构建1.2万栋带丰富元数据的精确多层建筑三维模型库

BuildingBRep-11K: Precise Multi-Storey B-Rep Building Solids with Rich Layout Metadata

  • 基于规则生成器自动构建符合设计规范的建筑体素
  • 可精准预测楼层数、房间数与平均面积,误差极低
  • 适合建筑生成、质量评估等方向研究者使用

随着人工智能发展,大规模三维建筑的自动生成成为热点,但训练仍需大量清洁且标注丰富的数据。本文提出BuildingBRep-11K,包含11,978栋2至10层的多层建筑(约10 GB),由遵循建筑学原则的形状语法管道生成。每个样本包含几何精确的B-rep实体(覆盖楼层、墙体、楼板及规则开口),并附带快速加载的.npy元数据文件,记录每层详细参数。生成过程融入空间尺度、采光优化与室内布局约束,并通过多阶段筛选(剔除布尔运算失败、过小房间与极端长宽比)确保符合建筑标准。为验证数据集可学习性,我们训练了两个轻量级PointNet基线:(i) 多属性回归——单编码器从4000点云预测楼层数、总房间数、每层向量与平均面积,在100个未见建筑上实现0.37层的MAE(87%在±1内)、5.7间房间的MAE、3.2 m²的平均面积MAE;(ii) 缺陷检测——同一骨干网络分类为良品或缺陷,在平衡的100模型集上达到54%准确率,召回82%真实缺陷,精度53%(41真阳性,9假阴性,37假阳性,13真阴性)。实验表明BuildingBRep-11K具备可学习性且对几何回归与拓扑质量评估均具挑战性。

原文摘要 · Abstract (English)

With the rise of artificial intelligence, the automatic generation of building-scale 3-D objects has become an active research topic, yet training such models still demands large, clean and richly annotated datasets. We introduce BuildingBRep-11K, a collection of 11 978 multi-storey (2-10 floors) buildings (about 10 GB) produced by a shape-grammar-driven pipeline that encodes established building-design principles. Every sample consists of a geometrically exact B-rep solid-covering floors, walls, slabs and rule-based openings-together with a fast-loading .npy metadata file that records detailed per-floor parameters. The generator incorporates constraints on spatial scale, daylight optimisation and interior layout, and the resulting objects pass multi-stage filters that remove Boolean failures, undersized rooms and extreme aspect ratios, ensuring compliance with architectural standards. To verify the dataset's learnability we trained two lightweight PointNet baselines. (i) Multi-attribute regression. A single encoder predicts storey count, total rooms, per-storey vector and mean room area from a 4 000-point cloud. On 100 unseen buildings it attains 0.37-storey MAE (87 \% within $\pm1$), 5.7-room MAE, and 3.2 m$^2$ MAE on mean area. (ii) Defect detection. With the same backbone we classify GOOD versus DEFECT; on a balanced 100-model set the network reaches 54 \% accuracy, recalling 82 \% of true defects at 53 \% precision (41 TP, 9 FN, 37 FP, 13 TN). These pilots show that BuildingBRep-11K is learnable yet non-trivial for both geometric regression and topological quality assessment

建筑生成B-rep数据集几何建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。