用扩散Transformer生成屋顶图,支持从草图到图像的多种重建。
Diffusion Transformers for Roof Graph Synthesis and Reconstruction

- 用扩散Transformer分两阶段生成屋顶顶点与连接关系。
- 图像引导重建时边准确率最高,优于现有方法。
- 可灵活切换条件,适用于草图、轮廓或影像输入。
我们提出RoofDiT,一种用于2D屋顶图合成与重建的生成框架。屋顶被紧凑地表示为角点与结构边构成的平面图,但现有方法常依赖固定几何规则或直接重建目标。RoofDiT将屋顶结构直接建模为顶点-边图,并学习其几何与连通性的条件生成先验。框架采用两阶段设计:扩散Transformer生成屋顶顶点,边预测模块推断对应图拓扑。为提升几何保真度,RoofDiT结合相对几何感知注意力与底面轮廓及航拍图像条件,同时使用对齐正则化鼓励常见的水平、垂直和对角屋顶模式。同一模型可通过改变条件信号支持无条件生成、底面轮廓条件合成和图像引导重建。实验表明,相比扩散基线,图生成质量更高;在底面轮廓条件下优于直线骨架先验;在图像引导重建中取得最高边F1分数。
原文摘要 · Abstract (English)
We present RoofDiT, a generative framework for 2D roof graph synthesis and reconstruction. Roofs are compactly described as planar graphs of junctions and structural edges, but existing methods often rely on fixed geometric rules or direct reconstruction objectives. RoofDiT instead models roof structures directly as vertex-edge graphs and learns a conditional generative prior over their geometry and connectivity. Our framework follows a two-stage design: a diffusion transformer generates roof vertices, and an edge prediction module infers the corresponding graph topology. To improve geometric fidelity, RoofDiT combines relative geometry-aware attention with footprint and aerial-image conditioning, while using an alignment regularizer to encourage common horizontal, vertical, and diagonal roof patterns. The same model supports unconditional generation, footprint-conditioned synthesis, and image-guided reconstruction by changing the conditioning signal. Experiments show improved graph generation quality over a diffusion baseline, favorable performance against a straight-skeleton prior in the footprint-conditioned setting, and the highest edge F1 among compared methods for image-guided reconstruction.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。