分析文本生成图像模型对残障群体的呈现,发现存在显著偏见。
Investigating Disability Representations in Text-to-Image Models
- 用结构化提示词对比通用与具体残障类别生成图像
- 发现不同残障类别的图像相似度高,体现代表性失衡
- 结合自动与人工评估,揭示情感基调负面倾向
文本到图像生成模型在高质量视觉内容生成方面取得显著进展,但其对社会群体的呈现仍存隐忧。尽管性别、种族等特征日益受到关注,残障群体的表征仍被严重忽视。本研究通过结构化提示设计,分析Stable Diffusion XL和DALL-E 3生成图像中残障人士的表现。通过比较通用残障提示与特定残障类别提示生成图像的相似性,评估表征差异。同时,采用情感极性分析,结合自动与人工评价,评估缓解策略对情感框架的影响。结果表明,残障表征存在持续性的不平衡,凸显出需持续评估与优化生成模型,以实现更具多样性和包容性的残障呈现。
原文摘要 · Abstract (English)
Text-to-image generative models have made remarkable progress in producing high-quality visual content from textual descriptions, yet concerns remain about how they represent social groups. While characteristics like gender and race have received increasing attention, disability representations remain underexplored. This study investigates how people with disabilities are represented in AI-generated images by analyzing outputs from Stable Diffusion XL and DALL-E 3 using a structured prompt design. We analyze disability representations by comparing image similarities between generic disability prompts and prompts referring to specific disability categories. Moreover, we evaluate how mitigation strategies influence disability portrayals, with a focus on assessing affective framing through sentiment polarity analysis, combining both automatic and human evaluation. Our findings reveal persistent representational imbalances and highlight the need for continuous evaluation and refinement of generative models to foster more diverse and inclusive portrayals of disability.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。