arXiv:2605.13113cs.CYcs.AI2026-05

为文生图模型的性别偏见审计提供场景化框架,让评估更贴合实际应用风险。

Context Matters: Auditing Gender Bias in T2I Generation through Risk-Tiered Use-Case Profiles

论文配图:Context Matters: Auditing Gender Bias in T2I Generation through Risk-Tiered Use-Case Profiles
图 1 · 摘自论文原文
  • 按欧盟AI法案风险等级设计使用场景分类,区分不同部署环境的审计要求。
  • 整合三类评测指标:性别预测、嵌入相似性与下游任务表现,系统量化偏见。
  • 针对不同场景提出具体危害类型,指导技术审计与政策制定。

文生图(T2I)生成模型正被广泛应用于教育、媒体和公共传播,逐步融入高影响力流程。由于生成图像常强化刻板印象,导致代表性缺失并影响人们对角色归属的认知,现有研究已提出多种量化性别偏见的指标。然而,当前评估仍碎片化:指标缺乏统一解释框架,未说明其测量目标、假设前提或在不同部署情境下的解读方式,限制了其在技术审计与治理讨论中的实用性。本文提出一个风险对齐的审计框架,包含三个核心部分:首先,依据欧盟AI法案的风险类别,构建风险分级的使用场景画像,说明为何审计期望应随部署情境和利益相关者暴露程度而变化;其次,建立指标目录,整合性别偏见评估方法,并按三大类别组织:性别预测、嵌入相似性、下游任务表现;第三,提出危害类型学,将上下文相关的危害(如代表性缺失、服务质量下降)映射到特定风险场景。最后,引入THUMB卡片(文本到图像危害感知使用场景对齐偏见度量),通过融合上下文、场景、偏见表现、危害假设与审计策略,实现系统性审计。

原文摘要 · Abstract (English)

Text-to-image (T2I) generative models are increasingly used to produce content for education, media, and public-facing communication, and are starting to be integrated into higher-impact pipelines. Since generated images tend to reinforce stereotypes, producing representational erasure via "default" depictions and shaping perceptions of who belongs in certain roles, a growing body of work has proposed metrics to quantify gender bias in T2I outputs. Yet existing evaluations remain fragmented. Metrics are often reported without a shared view of what they measure, what assumptions they entail, or how their results should be interpreted under different deployment contexts. This limits the usefulness of gender bias measurement for both technical auditing and emerging governance discussions. We propose a risk-aligned auditing framework for gender bias in T2I models composed of three constituents that connects risk categories, evaluation metrics, and harms. First, we identify risk-tiered use-case profiles aligned with the EU AI Act's risk categories to motivate why auditing expectations may vary with deployment contexts and stakeholder exposure. Second, we construct a metric catalog that consolidates gender-bias evaluation methods and organizes them in three measurement categories: gender prediction, embedding similarity, and downstream task. Third, we introduce a harm typology that maps context-dependent harm categories (e.g., representational, quality-of-service) to specific risk-tired scenarios. Finally, we introduce THUMB cards (Text-to-image Harms-informed Use-case-aligned Metrics of gender Bias) that help formulate auditing systematically by the incorporation of context, scenario and bias manifestation, harm hypotheses, and audit strategy.

文生图性别偏见风险评估AI治理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。