arXiv:2602.09775cs.CV2026-02

用大模型分析图文对定位数据来源,发现欧美占近半数,非洲南美严重不足。

Where Do Images Come From? Analyzing Captions to Geographically Profile Datasets

  • 通过LLM从标题提取地理信息,对图像数据集做国家级画像
  • 美英加三国占48%样本,非洲仅3.8%,南美1.8%
  • 高收入国家更易被收录,但数量多不等于多样性高

近期研究显示,文生图模型常无法生成具有地理代表性的图像,引发对其训练数据代表性的担忧。本文通过使用大语言模型从图文对中提取位置信息,对大规模多模态数据集进行地理画像。以三个常用数据集(Re-LAION、DataComp1B、Conceptual Captions)的英文标题为基础,针对20类常见实体(如house、flag),发现美国、英国和加拿大共占48.0%的样本,而南美洲和非洲国家分别仅有1.8%和3.8%的图像。观察到国家人均GDP与数据集代表性高度相关(ρ=0.82)。在对Re-LAION中4种非英语子集分析后发现,代表性显著偏向这些语言主要使用国。此外,高覆盖率并不意味着更高的视觉或语义多样性。最后,分析基于Re-LAION训练的Stable Diffusion v1.3生成的国家特定图像,发现虽外观真实,但覆盖范围远不及真实世界图像。

原文摘要 · Abstract (English)

Recent studies show that text-to-image models often fail to generate geographically representative images, raising concerns about the representativeness of their training data and motivating the question: which parts of the world do these training examples come from? We geographically profile large-scale multimodal datasets by mapping image-caption pairs to countries based on location information extracted from captions using LLMs. Studying English captions from three widely used datasets (Re-LAION, DataComp1B, and Conceptual Captions) across $20$ common entities (e.g., house, flag), we find that the United States, the United Kingdom, and Canada account for $48.0\%$ of samples, while South American and African countries are severely under-represented with only $1.8\%$ and $3.8\%$ of images, respectively. We observe a strong correlation between a country's GDP and its representation in the data ($ρ= 0.82$). Examining non-English subsets for $4$ languages from the Re-LAION dataset, we find that representation skews heavily toward countries where these languages are predominantly spoken. Additionally, we find that higher representation does not necessarily translate to greater visual or semantic diversity. Finally, analyzing country-specific images generated by Stable Diffusion v1.3 trained on Re-LAION, we show that while generations appear realistic, they are severely limited in their coverage compared to real-world images.

数据偏见地理画像多模态语言分布

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。