arXiv:2506.16898cs.AIcs.CV2025-06被引 2

扩散模型生成城市图像时隐含地理结构,但对美国多样地貌呈现刻板印象。

AI's Blind Spots: Geographic Knowledge and Diversity Deficit in Generated Urban Scenario

  • 用扩散模型生成各州街景图,通过嵌入特征对比地理分布
  • 美国通用提示导致沙漠、乡村等场景被低估,城市化特征过度集中
  • 揭示生成模型在城市规划中的地理偏见,适合关注AI公平性的研究者

基于扩散的文本到图像模型在城市分析与情景生成中日益普及,但其地理知识与表征偏差尚不明确。我们评估了FLUX 1-schnell和Stable Diffusion 3.5-Large在美国的表现,分别为每个州、州首府及通用“USA”提示生成150张街景图像。利用DINO-v2 ViT-S/14对图像进行嵌入,并通过Fréchet Inception Distance(FID)进行比较。成对FID聚类显示,地理邻近的州与首府常在空间上聚集,表明模型隐含地理结构。然而,通用“USA”提示使这种多样性坍缩为都市化刻板印象:边疆、沙漠、热带、乡村与小城环境在FID空间中显著被低估或远离主流分布。结果表明,尽管扩散模型能编码精细地理信息,仍会复现狭隘的国家级视觉偏见。

原文摘要 · Abstract (English)

Diffusion-based text-to-image models are increasingly used for urban analysis and scenario generation, but their geographic knowledge and representational biases remain poorly understood. We evaluate FLUX 1-schnell and Stable Diffusion 3.5-Large in the United States by generating 150 street-view images for each state, each state capital, and a generic "USA" prompt. Images are embedded with DINO-v2 ViT-S/14 and compared with Fréchet Inception Distance (FID). Pairwise FID clustering shows that geographically proximate states and capitals often group together, indicating implicit geographic structure. However, the generic ``USA'' prompt collapses this diversity into a metropolitan stereotype: frontier, desert, tropical, rural, and small-city environments are underrepresented or distant in FID space. These results show that diffusion models can encode fine-grained geography while still reproducing narrow national-scale visual stereotypes.

扩散模型地理偏见城市生成图像评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。