用大模型分析建筑风格,让跨文化比较更客观。
ArchiLense: A Framework for Quantitative Analysis of Architectural Styles Based on Vision Large Language Models
- 基于视觉语言模型构建建筑风格分析框架ArchiLense。
- 在1765张图上实现92.4%专家一致性与84.5%分类准确率。
- 适合做建筑文化比较、跨区域风格研究的学者和设计师。
不同地区的建筑文化具有风格多样性,受历史、社会、技术及地理条件影响。传统研究依赖主观专家判断和文献回顾,常存在地域偏见且解释范围有限。为此,本文提出三项贡献:(1)构建名为ArchDiffBench的专业建筑风格数据集,包含1,765张高质量建筑图像及其风格标注,覆盖不同地区与历史时期;(2)提出ArchiLense框架,基于视觉-语言模型,融合计算机视觉、深度学习与机器学习算法,实现建筑图像的自动识别、比较与精确分类,并生成描述性语言输出以阐明风格差异;(3)大量评估表明,ArchiLense在建筑风格识别中表现优异,与专家标注的一致率达92.4%,分类准确率为84.5%,有效捕捉图像间的风格差异。该方法克服了传统分析的主观性,为建筑文化比较研究提供了更客观、精准的新视角。
原文摘要 · Abstract (English)
Architectural cultures across regions are characterized by stylistic diversity, shaped by historical, social, and technological contexts in addition to geograph-ical conditions. Understanding architectural styles requires the ability to describe and analyze the stylistic features of different architects from various regions through visual observations of architectural imagery. However, traditional studies of architectural culture have largely relied on subjective expert interpretations and historical literature reviews, often suffering from regional biases and limited ex-planatory scope. To address these challenges, this study proposes three core contributions: (1) We construct a professional architectural style dataset named ArchDiffBench, which comprises 1,765 high-quality architectural images and their corresponding style annotations, collected from different regions and historical periods. (2) We propose ArchiLense, an analytical framework grounded in Vision-Language Models and constructed using the ArchDiffBench dataset. By integrating ad-vanced computer vision techniques, deep learning, and machine learning algo-rithms, ArchiLense enables automatic recognition, comparison, and precise classi-fication of architectural imagery, producing descriptive language outputs that ar-ticulate stylistic differences. (3) Extensive evaluations show that ArchiLense achieves strong performance in architectural style recognition, with a 92.4% con-sistency rate with expert annotations and 84.5% classification accuracy, effec-tively capturing stylistic distinctions across images. The proposed approach transcends the subjectivity inherent in traditional analyses and offers a more objective and accurate perspective for comparative studies of architectural culture.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。