用文本语义指导图像分割,自动识别建筑立面墙窗
Segment Any Architectural Facades (SAAF):An automatic segmentation model for building facades, walls and windows based on multimodal semantics guidance
- 融合文本与图像语义,提升对建筑构件的理解能力
- 端到端训练使模型在多数据集上达到更高分割精度
- 适合建筑视觉与多模态学习研究者参考
在建筑数字化背景下,自动分割墙体与窗户是提升建筑信息模型与计算机辅助设计效率的关键步骤。本文提出基于多模态语义引导的建筑立面自动分割模型SAAF(Segment Any Architectural Facades)。首先,SAAF采用多模态语义协同特征提取机制,结合自然语言处理技术,将文本描述中的语义信息与图像特征融合,增强对建筑立面组件的语义理解。其次,构建端到端训练框架,使模型能自主学习从文本描述到图像分割的映射关系,减少人工干预,提升分割自动化与鲁棒性。最后,在多个立面数据集上进行广泛实验,SAAF在mIoU指标上优于现有方法,表明其在多样化数据集下仍保持高精度分割能力。本模型在墙体与窗户分割任务的准确率与泛化能力上取得进展,为建筑计算机视觉技术发展提供参考,并探索了多模态学习在建筑领域的应用新路径。
原文摘要 · Abstract (English)
In the context of the digital development of architecture, the automatic segmentation of walls and windows is a key step in improving the efficiency of building information models and computer-aided design. This study proposes an automatic segmentation model for building facade walls and windows based on multimodal semantic guidance, called Segment Any Architectural Facades (SAAF). First, SAAF has a multimodal semantic collaborative feature extraction mechanism. By combining natural language processing technology, it can fuse the semantic information in text descriptions with image features, enhancing the semantic understanding of building facade components. Second, we developed an end-to-end training framework that enables the model to autonomously learn the mapping relationship from text descriptions to image segmentation, reducing the influence of manual intervention on the segmentation results and improving the automation and robustness of the model. Finally, we conducted extensive experiments on multiple facade datasets. The segmentation results of SAAF outperformed existing methods in the mIoU metric, indicating that the SAAF model can maintain high-precision segmentation ability when faced with diverse datasets. Our model has made certain progress in improving the accuracy and generalization ability of the wall and window segmentation task. It is expected to provide a reference for the development of architectural computer vision technology and also explore new ideas and technical paths for the application of multimodal learning in the architectural field.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。