arXiv:2604.21036cs.AI2026-04被引 1

用户可自定义公平性目标,用提示词调控生成图像的族裔分布。

Who Defines Fairness? Target-Based Prompting for Demographic Representation in Generative Models

论文配图:Who Defines Fairness? Target-Based Prompting for Demographic Representation in Generative Models
图 1 · 摘自论文原文
  • 通过提示词改写实现推理时公平性干预,无需重训练模型。
  • 36个提示词实验中,肤色分布向设定目标偏移,偏差显著降低。
  • 支持多种公平性定义,适合关注生成内容社会影响的使用者。

Stable Diffusion和DALL-E等文本到图像(T2I)模型虽已普及,但常复制社会偏见,例如“医生”或“首席执行官”等职业提示多生成浅肤色人物,而“清洁工”等低地位角色则更多样,强化刻板印象。现有缓解方法多需重新训练或使用定制数据集,对普通用户不友好。本文提出一种轻量级、推理时干预的框架,通过提示词层面调整实现代表性偏差缓解,无需修改底层模型。不预设单一公平性标准,允许用户在多种定义间选择——从均匀分布到由大语言模型(LLM)支持、带来源引用与置信度估计的复杂定义。这些分布指导生成按比例匹配的族裔特定提示变体,并通过审计声明目标的符合度及实际肤色分布来评估效果,而非默认均匀即公平。在涵盖30种职业与6种非职业场景的36个提示测试中,本方法使观察到的肤色结果朝声明目标方向转移,并在目标直接以肤色空间定义时(备用方案),显著减少偏离。该工作展示了公平性干预如何在推理阶段透明、可控且可用,直接赋能生成式AI用户。

原文摘要 · Abstract (English)

Text-to-image(T2I) models like Stable Diffusion and DALL-E have made generative AI widely accessible, yet recent studies reveal that these systems often replicate societal biases, particularly in how they depict demographic groups across professions. Prompts such as 'doctor' or 'CEO' frequently yield lighter-skinned outputs, while lower-status roles like 'janitor' show more diversity, reinforcing stereotypes. Existing mitigation methods typically require retraining or curated datasets, making them inaccessible to most users. We propose a lightweight, inference-time framework that mitigates representational bias through prompt-level intervention without modifying the underlying model. Instead of assuming a single definition of fairness, our approach allows users to select among multiple fairness specifications-ranging from simple choices such as a uniform distribution to more complex definitions informed by a large language model(LLM) that cites sources and provides confidence estimates. These distributions guide the construction of demographic specific prompt variants in the corresponding proportions, and we evaluate alignment by auditing adherence to the declared target and measuring the resulting skin tone distribution rather than assuming uniformity as 'fairness'. Across 36 prompts spanning 30 occupations and 6 non-occupational contexts, our method shifts observed skin-tone outcomes in directions consistent with the declared target, and reduces deviation from targets when the target is defined directly in skin-tone space(fallback). This work demonstrates how fairness interventions can be made transparent, controllable, and usable at inference time, directly empowering users of generative AI.

生成模型公平性提示工程肤色偏见

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。