arXiv:2608.29992cs.CV2026-08中稿 · publication in the…

用AI从街景图自动重建3D城市建筑立面细节,无需大量标注数据。

SVI2LoD3: Agent-Driven Reconstruction of LoD3 Facade Openings in Semantic 3D City Models from Volunteered Street View Imagery using Large Language and Visual Models

论文配图:SVI2LoD3: Agent-Driven Reconstruction of LoD3 Facade Openings in Semantic 3D City Models from Volunteered Street View Imagery using Large Language and Visual Models
图 1 · 摘自论文原文
  • 通过智能代理驱动的端到端流程,实现零样本分割重建。
  • 在eTRIMS数据集上达到高精度,且输出符合CityGML标准。
  • 提出新评估指标FFD,更准确衡量立面语义与布局质量。

本文提出一种端到端、代理驱动的流水线,用于从志愿者街景影像中重建3D城市模型的LoD3立面开口,直接生成符合CityGML规范的输出。与依赖大量人工标注训练数据的监督语义分割方法不同,该方法采用零样本分割策略,显著降低标注成本,同时在eTRIMS数据集基准测试中仍保持优异性能。另一关键贡献是强制保证部件层次结构正确性,从而生成符合CityGML规范的LoD3建筑模型。此外,本文提出一种新型立面重建评估指标——立面特征距离(Facade Feature Distance, FFD)。不同于传统的像素级重叠度量如mIoU或FRDS,FFD基于视觉变换器提取的高层特征空间计算距离,能同时捕捉语义准确性和建筑布局合理性,提供更适配立面重建质量的评估方式。所提流水线与评估策略共同为语义增强型3D城市模型的自动化生成与分析提供了实用且可扩展的解决方案。代码已公开于:https://github.com/hcu-cml/citydb-SVI2LoD3-ai。

原文摘要 · Abstract (English)

This paper presents an end-to-end, agent-driven pipeline for the LoD3 reconstruction of facade openings in 3D city models, producing directly usable CityGML-conform outputs. In contrast to existing approaches that rely on supervised semantic segmentation and therefore require large amounts of manually annotated training data, the proposed method employs a zero-shot segmentation strategy. This substantially reduces the annotation effort while still achieving strong performance in our benchmark on the eTRIMS dataset. A further key contribution is the enforcement of correct partonomic hierarchies, thereby producing CityGML-conform LoD3 building models. Beyond the reconstruction pipeline itself, this work also introduces a novel evaluation metric for facade reconstruction, termed Facade Feature Distance (FFD). Unlike conventional metrics such as mIoU or FRDS, which assess similarity primarily through pixel-wise overlap, FFD measures distance in a high-level feature space derived from a vision transformer. In doing so, it captures both semantic correctness and architectural layout, providing a more suitable assessment of facade reconstruction quality. The proposed pipeline and evaluation strategy together offer a practical and scalable contribution toward the automated generation and analysis of semantically enriched 3D city models. The developed code is published at: https://github.com/hcu-cml/citydb-SVI2LoD3-ai.

3D城市建模立面重建零样本学习评估指标

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。