arXiv:2505.01430cs.CV2025-05被引 7

提出新方法诊断文生图模型的文化偏见,发现非西方文化生成效果显著更差。

Deconstructing Bias: A Multifaceted Framework for Diagnosing Cultural and Compositional Inequities in Text-to-Image Generative Models

  • 用组件包含度评分(CIS)量化图像生成中文化语境的忠实度。
  • 2400张图像分析显示,非西方文化提示词生成质量明显低于西方文化。
  • 揭示数据不平衡与注意力熵是导致偏见的关键因素,适合关注AI公平性的研究者参考。

文生图(T2I)模型的变革潜力依赖于其从文本提示生成具有文化多样性和逼真感图像的能力。然而,这些模型常会复制训练数据中的文化偏见,导致系统性误表。本文引入组件包含度评分(CIS),用于评估不同文化背景下图像生成的保真度。通过分析2400张图像,我们量化了组合脆弱性和上下文错位问题,揭示了西方与非西方文化提示间显著的性能差距。研究指出数据失衡、注意力熵和嵌入叠加是影响模型公平性的关键因素。通过对Stable Diffusion等模型进行CIS基准测试,本文为提升AI生成图像的文化包容性提供了架构与数据层面的干预思路。本工作推动了对文生图模型偏见的诊断与缓解,倡导构建更具公平性的AI系统。

原文摘要 · Abstract (English)

The transformative potential of text-to-image (T2I) models hinges on their ability to synthesize culturally diverse, photorealistic images from textual prompts. However, these models often perpetuate cultural biases embedded within their training data, leading to systemic misrepresentations. This paper benchmarks the Component Inclusion Score (CIS), a metric designed to evaluate the fidelity of image generation across cultural contexts. Through extensive analysis involving 2,400 images, we quantify biases in terms of compositional fragility and contextual misalignment, revealing significant performance gaps between Western and non-Western cultural prompts. Our findings underscore the impact of data imbalance, attention entropy, and embedding superposition on model fairness. By benchmarking models like Stable Diffusion with CIS, we provide insights into architectural and data-centric interventions for enhancing cultural inclusivity in AI-generated imagery. This work advances the field by offering a comprehensive tool for diagnosing and mitigating biases in T2I generation, advocating for more equitable AI systems.

文生图文化偏见公平性评估指标

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。