arXiv:2504.06313cs.CVcs.CY2025-04

研究主流文生图模型如何表现不同国籍人群的日常活动形象

Text-to-Image Models and Their Representation of People from Different Nationalities Engaging in Activities

  • 用206个国家别+5类活动生成2060张图,分析视觉呈现模式
  • 近三成图像出现不切实际的传统服饰,中东、撒哈拉以南非洲更严重
  • 模型常自动添加'传统'一词,导致图文匹配度虚高

本文研究主流文生图模型DALL-E 3与Gemini 3 Pro Preview在生成206个国籍人群参与五类日常活动图像时的表现。共设计五个场景,生成2,060张图像。跨活动与模型统计显示,28.4%的图像中人物穿着传统服饰,其中多例服饰明显不适用于所指定活动。该现象与地理区域显著相关,中东与北非、撒哈拉以南非洲地区尤为突出,并与世界银行收入组别有关。在两项体育活动中,类似区域与收入关联的不实着装模式亦被观察到。通过CLIP、ALIGN与GPT-4.1 mini对9,270个图像-提示对进行图文对齐评估,发现标注为传统服饰的图像在提示中包含国名时获得更高对齐评分,而移除国名后该趋势减弱或逆转。进一步提示分析表明,某一模型在传统服饰图像中频繁添加‘traditional’一词(50.3%),而在其他图像中仅16.6%出现。结果表明,这些表征模式受生成器、评估模型及提示修正等多环节共同影响。

原文摘要 · Abstract (English)

This paper investigates how popular text-to-image (T2I) models, DALL-E 3 and Gemini 3 Pro Preview, depict people from 206 nationalities when prompted to generate images of individuals engaging in common everyday activities. Five scenarios were developed, and 2,060 images were generated using input prompts that specified nationalities across five activities. When aggregating across activities and models, results showed that 28.4% of the images depicted individuals wearing traditional attire, including attire that is impractical for the specified activities in several cases. This pattern was statistically significantly associated with regions, with the Middle East & North Africa and Sub-Saharan Africa disproportionately affected, and was also associated with World Bank income groups. Similar region- and income-linked patterns were observed for images labeled as depicting impractical attire in two athletics-related activities. To assess image-text alignment, CLIP, ALIGN, and GPT-4.1 mini were used to score 9,270 image-prompt pairs. Images labeled as featuring traditional attire received statistically significantly higher alignment scores when prompts included country names, and this pattern weakened or reversed when country names were removed. Revised prompt analysis showed that one model frequently inserted the word "traditional" (50.3% for traditional-labeled images vs. 16.6% otherwise). These results indicate that these representational patterns can be shaped by several components of the pipeline, including image generator, evaluation models, and prompt revision.

文生图偏见检测图像生成文化表征

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。