揭示艺术风格中隐藏的性别偏见,并发现生成模型会放大这些偏见。
Gender Artifacts from Art History to Text-to-Image Generation

- 构建首个跨历史与生成图像的性别-风格数据集,支持直接对比。
- 发现不同艺术风格中性别特征通过笔触、色彩等视觉元素被编码。
- 生成模型在文本到图像生成中会放大原始艺术中的性别刻板印象。
艺术风格根植于特定的社会历史背景,蕴含社会等级关系,包括对性别的独特建构。然而,在人工智能研究中,风格长期被视为表面视觉属性:仅影响颜色、笔触和纹理的滤镜,作用于内容中立的场景。本文提出首个研究历史与生成图像中性别表现与风格相互作用的数据集——StyleGender,包含7.4万张图像,覆盖19种艺术风格,涵盖有风格与性别标注的历史艺术图像、在受控风格与性别提示下生成的文本到图像(T2I)图像,以及语义对齐的图像集,支持艺术史到生成结果的直接比较。我们提出两种性别偏差度量(SGA):PixelSGA与MaskSGA,分别捕捉像素级与构图结构中的性别信号。结果显示:(1) 性别表征在多种艺术风格中塑造了特定视觉特征;(2) 风格关键词将此类模式引入到文本到图像生成中;(3) 生成模型倾向于在输出中放大历史源中已存在的性别偏差。
原文摘要 · Abstract (English)
Artistic styles are rooted in specific socio-historical contexts that encode social hierarchies, including distinct constructions of gender. Yet in AI research, style has long been treated as a surface-level visual property: a filter of color, brushstroke, and texture applied to otherwise content-neutral scenes. We introduce the first dataset to investigate the interplay between gender representation and style in both historical and generated images. StyleGender comprises 74k images spanning 19 artistic styles, comprising art historical images with style and gender annotations, T2I-generated images under controlled style and gender prompts, and a semantically aligned set enabling direct art history-to-generation comparison. By proposing two Set Gender Artifact (SGA) metrics (PixelSGA and MaskSGA), capturing gender signals at the pixel level and in compositional structure, we show that (1) gender representation shapes visual features across artistic styles, (2) style keywords carry these patterns into T2I generation, and (3) generative models tend to amplify gender artifacts beyond what is observed in historical sources.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。