arXiv:2604.09923cs.AIcs.CV2026-04

用生成画像揭示AI图像模型的偏见,让普通人一眼看懂。

GLEaN: A Text-to-image Bias Detection Approach for Public Comprehension

  • 通过生成大量身份图像并合成代表脸,直观呈现模型偏见
  • 在291名用户中验证,可视化效果等同表格但耗时更短
  • 无需模型内部数据,适用于任何闭源系统

文本到图像(T2I)模型及其隐含偏见正日益影响公众接触的视觉内容。尽管已有大量研究关注T2I系统的偏见测量、审计与缓解,但这些方法主要面向技术专家,缺乏公众可读性。我们提出GLEaN(多尺度生成相似度评估),一种基于肖像的可解释性流程,使T2I模型偏见对普通大众直观可见。GLEaN包含三个阶段:从身份提示自动生成大规模图像、基于面部关键点的过滤与空间对齐、以及中位像素合成,将模型的典型倾向浓缩为一张代表性肖像。生成的合成图无需统计背景即可理解——观众一眼就能看出模型在提示‘医生’与‘罪犯’时分别‘想象’出谁。我们在Stable Diffusion XL上对40个社会与职业身份提示进行了测试,生成的画像复现了已知偏见,并揭示了肤色与预测情绪之间的新关联。一项含291名受试者的组间用户研究显示,GLEaN肖像在传达偏见方面与传统数据表效果相当,但所需观看时间显著更短。由于该方法仅依赖生成结果,可应用于任意黑盒或闭源系统,无需访问模型内部。GLEaN提供了一种可扩展、模型无关的偏见可解释性方案,专为公众理解设计,代码已开源:https://github.com/cultureiolab/GLEaN。

原文摘要 · Abstract (English)

Text-to-image (T2I) models, and their encoded biases, increasingly shape the visual media the public encounters. While researchers have produced a rich body of work on bias measurement, auditing, and mitigation in T2I systems, those methods largely target technical stakeholders, leaving a gap in public legibility. We introduce GLEaN (Generative Likeness Evaluation at N-Scale), a portrait-based explainability pipeline designed to make T2I model biases visually understandable to a broad audience. GLEaN comprises three stages: automated large-scale image generation from identity prompts, facial landmark-based filtering and spatial alignment, and median-pixel composition that distills a model's central tendency into a single representative portrait. The resulting composites require no statistical background to interpret; a viewer can see, at a glance, who a model 'imagines' when prompted with 'a doctor' versus a 'felon.' We demonstrate GLEaN on Stable Diffusion XL across 40 social and occupational identity prompts, producing composites that reproduce documented biases and surface new associations between skin tone and predicted emotion. We find in a between-subjects user study (N = 291) that GLEaN portraits communicate biases as effectively as conventional data tables, but require significantly less viewing time. Because the method relies solely on generated outputs, it can also be replicated on any black-box and closed-weight systems without access to model internals. GLEaN offers a scalable, model-agnostic approach to bias explainability, purpose-built for public comprehension, and is publicly available at https://github.com/cultureiolab/GLEaN.

AI偏见可解释性图像生成公众理解

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。