探究文本生成图像中的社会偏见在叙事图像中如何加剧
Investigating Social Bias in Narrative Image Generation

- 扩展评估框架至分镜与漫画,对比不同视觉形式的偏见表现
- 专有模型在漫画生成中偏见率比照片高18.2个百分点
- 叙事结构和文字元素使偏见更明显,适合内容安全研究者阅读
文本到图像(T2I)生成模型正广泛应用于媒体创作与教育领域,引发对其输出可能再现社会偏见的担忧。现有研究多聚焦于照片生成任务,但尚不清楚此类偏见在分镜、漫画等叙事性视觉格式中如何表现。本文通过将BBG文本偏见评估框架适配至图像生成,对比六种T2I模型在照片、分镜和漫画生成中的偏见表达。结果显示,专有模型在照片生成中平均存在25.9%的偏见输出,分镜生成中提升9.6个百分点,漫画生成中进一步上升18.2个百分点。分析发现,照片中的偏见主要通过细微视觉线索体现,而分镜与漫画则通过事件编排、角色位置、叙事结局及文字元素更明确地暴露偏见。这表明部分在照片中隐蔽的偏见在叙事视觉格式中会显著显现,凸显了超越照片评估T2I系统多样性的必要性。
原文摘要 · Abstract (English)
Text-to-image (T2I) generation models are increasingly embedded in applications such as media content creation and education, raising concerns about how their outputs may reproduce social biases. Prior work has shown that T2I models exhibit social biases, yet existing evaluations largely focus on a photo generation task. As a result, it remains unclear whether and how such biases manifest in more narrative visual formats, such as storyboards and comics, where characters and events are presented across multiple panels. In this work, we compare bias expression across photo, storyboard, and comic generation in six T2I models by adapting BBG, a text-based bias evaluation framework, to image generation. Our results show that proprietary models generate 25.9% biased outputs in photo generation on average, with biased outputs increasing by 9.6pp in storyboard generation and 18.2pp in comic generation. We also find that photos mainly encode biases through subtle visual cues, while storyboards and comics reveal them more explicitly through event sequencing, character positioning, narrative resolution, and textual elements. These findings show that biases that remain less visible in photo generation may surface in narrative visual formats, highlighting the importance of evaluating T2I systems with diverse visual formats beyond photo generation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。