arXiv:2511.13750cs.LGcs.AI2025-11中稿 · WACV2025被引 2

用自然语言自动挖掘扩散模型中的隐空间偏见,无需标注或重训练。

SCALEX: Scalable Concept and Latent Exploration for Diffusion Models

  • 仅靠自然语言提示从隐空间提取语义方向,实现零样本解析。
  • 发现职业相关性别偏见,识别身份描述的语义对齐程度。
  • 无需监督即可揭示概念聚类结构,适合大规模偏见分析。

图像生成模型常包含社会偏见,如与性别、种族和职业相关的刻板印象。现有分析方法或局限于预定义类别,或依赖人工解读隐空间方向,难以扩展且难发现细微或意外模式。我们提出SCALEX框架,实现扩散模型隐空间的可扩展、自动化探索。该方法仅通过自然语言提示从H空间提取语义有意义的方向,实现零样本解释,无需重训练或标注。这使得任意概念间可系统比较,支持大规模发现模型内部关联。实验表明,SCALEX能检测职业提示中的性别偏见,排序身份描述符的语义对齐度,并在无监督下揭示概念聚类结构。通过直接将提示与隐空间方向关联,相比以往方法,使扩散模型偏见分析更具可扩展性、可解释性和可拓展性。

原文摘要 · Abstract (English)

Image generation models frequently encode social biases, including stereotypes tied to gender, race, and profession. Existing methods for analyzing these biases in diffusion models either focus narrowly on predefined categories or depend on manual interpretation of latent directions. These constraints limit scalability and hinder the discovery of subtle or unanticipated patterns. We introduce SCALEX, a framework for scalable and automated exploration of diffusion model latent spaces. SCALEX extracts semantically meaningful directions from H-space using only natural language prompts, enabling zero-shot interpretation without retraining or labelling. This allows systematic comparison across arbitrary concepts and large-scale discovery of internal model associations. We show that SCALEX detects gender bias in profession prompts, ranks semantic alignment across identity descriptors, and reveals clustered conceptual structure without supervision. By linking prompts to latent directions directly, SCALEX makes bias analysis in diffusion models more scalable, interpretable, and extensible than prior approaches.

扩散模型偏见分析隐空间探索

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。