arXiv:2510.00045cs.CVcs.AI2025-10被引 3

六款文生图模型生成医护图像时普遍存在性别刻板印象。

Beyond the Prompt: Gender Bias in Text-to-Image Models, with a Case Study on Hospital Professions

  • 用精准提示词生成100张职业肖像图,系统测试模型性别倾向。
  • 护士全为女性,外科医生多为男性,不同模型差异显著。
  • 提示词如'企业风'强化男性形象,'美丽'偏好女性,需警惕误导。

文生图模型在专业、教育和创意场景中应用日益广泛,但其输出常隐含并放大社会偏见。本文研究六款先进开源模型(HunyuanImage 2.1、HiDream-I1-dev、Qwen-Image、FLUX.1-dev、Stable-Diffusion 3.5 Large、Stable-Diffusion-XL)在医院职业中的性别表现。通过精心设计的提示词,为五类职业(心脏病专家、医院主任、护士、急救员、外科医生)与五种肖像风格(无、企业风、中性、审美、美丽)组合生成各100张图像。分析显示:所有模型均将护士呈现为女性,外科医生以男性为主。模型间存在差异:Qwen-Image与SDXL强化男性主导,HiDream-I1-dev结果混杂,而FLUX.1-dev在多数角色中偏向女性。肖像风格进一步影响性别分布,如‘企业风’强化男性形象,‘美丽’更倾向女性。敏感度差异大:Qwen-Image几乎不受提示影响,而FLUX.1-dev、SDXL与SD3.5对提示高度敏感。结果表明,文生图模型的性别偏见既系统又模型特异。除揭示偏差外,本文强调提示词在塑造人物形象中的关键作用。呼吁采用抗偏见设计、平衡默认设置及用户引导,防止生成式AI固化职业刻板印象。

原文摘要 · Abstract (English)

Text-to-image (TTI) models are increasingly used in professional, educational, and creative contexts, yet their outputs often embed and amplify social biases. This paper investigates gender representation in six state-of-the-art open-weight models: HunyuanImage 2.1, HiDream-I1-dev, Qwen-Image, FLUX.1-dev, Stable-Diffusion 3.5 Large, and Stable-Diffusion-XL. Using carefully designed prompts, we generated 100 images for each combination of five hospital-related professions (cardiologist, hospital director, nurse, paramedic, surgeon) and five portrait qualifiers ("", corporate, neutral, aesthetic, beautiful). Our analysis reveals systematic occupational stereotypes: all models produced nurses exclusively as women and surgeons predominantly as men. However, differences emerge across models: Qwen-Image and SDXL enforce rigid male dominance, HiDream-I1-dev shows mixed outcomes, and FLUX.1-dev skews female in most roles. HunyuanImage 2.1 and Stable-Diffusion 3.5 Large also reproduce gender stereotypes but with varying degrees of sensitivity to prompt formulation. Portrait qualifiers further modulate gender balance, with terms like corporate reinforcing male depictions and beautiful favoring female ones. Sensitivity varies widely: Qwen-Image remains nearly unaffected, while FLUX.1-dev, SDXL, and SD3.5 show strong prompt dependence. These findings demonstrate that gender bias in TTI models is both systematic and model-specific. Beyond documenting disparities, we argue that prompt wording plays a critical role in shaping demographic outcomes. The results underscore the need for bias-aware design, balanced defaults, and user guidance to prevent the reinforcement of occupational stereotypes in generative AI.

文生图性别偏见医疗影像提示词

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。