用大模型理解建筑,让机器读懂古今建筑的美学与文化。
ArchGPT: Understanding the World's Architectures with Large Multimodal Models
- 构建了31.5万条建筑视觉问答数据,支持多模态理解。
- 通过3D重建与语义分割筛选高质量建筑图像,提升数据可靠性。
- 适合建筑研究、文化遗产保护与智能设计领域的学者与从业者。
建筑承载着审美、文化和历史价值,是人类文明的实体见证。研究人员长期利用虚拟现实(VR)、混合现实(MR)和增强现实(AR)技术实现对建筑的沉浸式探索与解读,提升了教育、遗产保护与专业设计中的可及性与理解度。然而,现有系统多为定制开发,依赖硬编码标注与特定任务交互,难以跨多样建筑环境扩展。本文提出ArchGPT,一个面向建筑领域的多模态视觉问答(VQA)模型,并构建了一套可扩展的数据构建流程,用于生成高质量、领域专精的建筑类VQA标注。该流程产出约31.5万条图像-问题-答案三元组的Arch-300K数据集。其构建过程包括:从Wikimedia Commons收集建筑场景,采用新颖的粗到精策略结合3D重建与语义分割,筛选无遮挡、结构一致的建筑图像;为降低原始文本元数据的噪声与不一致性,提出基于大语言模型的文本验证与知识蒸馏流程,生成可靠的建筑专属问答对;在此基础上,进一步合成形式分析注释——包括详细描述与基于维度的对话,以提供更丰富的语义多样性且保持数据忠实性。我们使用Arch-300K对开源多模态基座模型ShareGPT4V-7B进行监督微调,得到ArchGPT。
原文摘要 · Abstract (English)
Architecture embodies aesthetic, cultural, and historical values, standing as a tangible testament to human civilization. Researchers have long leveraged virtual reality (VR), mixed reality (MR), and augmented reality (AR) to enable immersive exploration and interpretation of architecture, enhancing accessibility, public understanding, and creative workflows around architecture in education, heritage preservation, and professional design practice. However, existing VR/MR/AR systems are often developed case-by-case, relying on hard-coded annotations and task-specific interactions that do not scale across diverse built environments. In this work, we present ArchGPT, a multimodal architectural visual question answering (VQA) model, together with a scalable data-construction pipeline for curating high-quality, architecture-specific VQA annotations. This pipeline yields Arch-300K, a domain-specialized dataset of approximately 315,000 image-question-answer triplets. Arch-300K is built via a multi-stage process: first, we curate architectural scenes from Wikimedia Commons and filter unconstrained tourist photo collections using a novel coarse-to-fine strategy that integrates 3D reconstruction and semantic segmentation to select occlusion-free, structurally consistent architectural images. To mitigate noise and inconsistency in raw textual metadata, we propose an LLM-guided text verification and knowledge-distillation pipeline to generate reliable, architecture-specific question-answer pairs. Using these curated images and refined metadata, we further synthesize formal analysis annotations-including detailed descriptions and aspect-guided conversations-to provide richer semantic variety while remaining faithful to the data. We perform supervised fine-tuning of an open-source multimodal backbone ,ShareGPT4V-7B, on Arch-300K, yielding ArchGPT.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。