AI绘图模型在沙特职业形象生成中强化性别刻板印象,女性比例严重偏低。
Gender Stereotypes in Professional Roles Among Saudis: An Analytical Study of AI-Generated Images Using Language Models
- 用中性提示词生成56种沙特职业图像,由本地专家评分
- DALL-E V3生成图像男性占比达96%,领导与技术岗位最严重
- 文化误读常被误认为反刻板表达,需更精准评估框架
本研究探讨当代文生图AI模型在生成沙特专业人员形象时,是否延续性别刻板印象与文化偏差。我们使用ImageFX、DALL-E V3和Grok对56种沙特职业生成1,006张图像,采用中性提示词。两名受训的沙特标注员对每张图像在五个维度(性别感知、着装外貌、背景环境、活动互动、年龄)进行评估,第三名资深研究员在分歧时裁决,共获得10,100个独立判断。结果显示性别失衡显著:ImageFX输出中85%为男性,Grok为86.6%,而DALL-E V3高达96%,表明其性别刻板印象最强。该现象在领导与技术岗位尤为突出。所有模型均频繁出现服饰、场景及活动的文化不准确问题。反刻板图像多源于文化误读而非真实进步。研究结论指出,当前模型反映训练数据中人类社会偏见,仅有限呈现沙特劳动力市场的性别结构与文化细节。亟需更多元训练数据、更公平算法与文化敏感评估体系,以实现公正且真实的视觉输出。
原文摘要 · Abstract (English)
This study investigates the extent to which contemporary Text-to-Image artificial intelligence (AI) models perpetuate gender stereotypes and cultural inaccuracies when generating depictions of professionals in Saudi Arabia. We analyzed 1,006 images produced by ImageFX, DALL-E V3, and Grok for 56 diverse Saudi professions using neutral prompts. Two trained Saudi annotators evaluated each image on five dimensions: perceived gender, clothing and appearance, background and setting, activities and interactions, and age. A third senior researcher adjudicated whenever the two primary raters disagreed, yielding 10,100 individual judgements. The results reveal a strong gender imbalance, with ImageFX outputs being 85\% male, Grok 86.6\% male, and DALL-E V3 96\% male, indicating that DALL-E V3 exhibited the strongest overall gender stereotyping. This imbalance was most evident in leadership and technical roles. Moreover, cultural inaccuracies in clothing, settings, and depicted activities were frequently observed across all three models. Counter-stereotypical images often arise from cultural misinterpretations rather than genuinely progressive portrayals. We conclude that current models mirror societal biases embedded in their training data, generated by humans, offering only a limited reflection of the Saudi labour market's gender dynamics and cultural nuances. These findings underscore the urgent need for more diverse training data, fairer algorithms, and culturally sensitive evaluation frameworks to ensure equitable and authentic visual outputs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。