评测五款AI图像生成器在建筑风格上的准确性,发现普遍不准且易出错。
Architecture inside the mirage: evaluating generative image models on architectural style, elements, and typologies
- 用30个建筑提示词测试5大平台,每组生成4图共600张
- 平均准确率仅42%,常见提示准确率是罕见提示的2.7倍
- 普遍存在过度装饰、风格混淆等问题,适合教育者谨慎使用
生成式人工智能(GenAI)文本到图像系统正被广泛用于生成建筑图像,但其在历史规则严格的建筑领域中生成准确图像的能力仍缺乏评估。我们使用30个涵盖风格、类型和规范元素的建筑提示词,评估了五个主流GenAI图像平台(Adobe Firefly、DALL-E 3、Google Imagen 3、Microsoft Image Generator和Midjourney)。每个提示-生成器组合产出4张图像(共600张)。两名建筑史专家独立评分,分歧由共识解决。以每组4张图中准确数量(0至4)汇总表现。常见提示生成的图像准确率比罕见提示高2.7倍(p < 0.05)。各平台整体准确率有限(最高52%,最低32%,均值42%)。全部正确(4/4)结果在各平台间相似;而全部错误(0/4)结果差异显著,Imagen 3失败最少,Microsoft Image Generator失败最多。对图像数据集的定性分析发现反复出现的问题:过度装饰、中世纪风格与其后现代复兴混淆、误读描述性提示(如蛋形线饰、带状柱、穹顶结构)。研究支持对GenAI合成内容进行可见标注,建立未来训练数据集的溯源标准,并审慎用于教育场景。
原文摘要 · Abstract (English)
Generative artificial intelligence (GenAI) text-to-image systems are increasingly used to generate architectural imagery, yet their capacity to reproduce accurate images in a historically rule-bound field remains poorly characterized. We evaluated five widely used GenAI image platforms (Adobe Firefly, DALL-E 3, Google Imagen 3, Microsoft Image Generator, and Midjourney) using 30 architectural prompts spanning styles, typologies, and codified elements. Each prompt-generator pair produced four images (n = 600 images total). Two architectural historians independently scored each image for accuracy against predefined criteria, resolving disagreements by consensus. Set-level performance was summarized as zero to four accurate images per four-image set. Image output from Common prompts was 2.7-fold more accurate than from Rare prompts (p < 0.05). Across platforms, overall accuracy was limited (highest accuracy score 52 percent; lowest 32 percent; mean 42 percent). All-correct (4 out of 4) outcomes were similar across platforms. By contrast, all-incorrect (0 out of 4) outcomes varied substantially, with Imagen 3 exhibiting the fewest failures and Microsoft Image Generator exhibiting the highest number of failures. Qualitative review of the image dataset identified recurring patterns including over-embellishment, confusion between medieval styles and their later revivals, and misrepresentation of descriptive prompts (for example, egg-and-dart, banded column, pendentive). These findings support the need for visible labeling of GenAI synthetic content, provenance standards for future training datasets, and cautious educational use of GenAI architectural imagery.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。