arXiv:2510.20042cs.CV2025-10中稿 · IASEAI 2026被引 6

评测生成图像模型的文化偏见,发现南北文化差异被弱化。

Exposing Blindspots: Cultural Bias Evaluation in Generative Image Models

  • 构建跨国家、跨时代、多类别评估框架,统一测试图文生成与图像编辑
  • 发现模型默认生成全球北方现代风格,忽略文化差异;编辑过程更削弱文化真实感
  • 揭示当前模型仅做表面修饰,缺乏语境一致的深层文化适配,适合研究公平性者参考

生成式图像模型虽能产出高质量视觉内容,却常出现文化误表。以往研究主要聚焦文本到图像(T2I)系统中的文化偏见,而图像到图像(I2I)编辑器则未受充分关注。本文通过覆盖六个国家、8个类别36个子类别的评估体系,结合时代感知提示,在标准化协议下同时审计T2I生成与I2I编辑。采用固定设置的开源模型,进行跨国家、跨时代、跨类别评估。方法融合标准自动指标、文化敏感的检索增强型VQA及母语专家的人工判断。为确保可复现性,完整公开图像数据集、提示词与配置。研究发现:(1) 在无国家特定提示下,模型倾向于输出以全球北方为主、偏现代的统一图示,抹平国别差异;(2) 迭代式I2I编辑会削弱文化真实性,即便传统指标不变或改善;(3) I2I模型仅使用色调变化、通用道具等表面线索,而非符合时代背景的上下文一致调整,对全球南方目标常保留源身份。结果表明当前系统在文化敏感编辑上仍不可靠。通过发布标准化数据、提示与人工评估协议,我们提供了一个可复现、以文化为中心的基准,用于诊断和追踪生成图像模型中的文化偏见。

原文摘要 · Abstract (English)

Generative image models produce striking visuals yet often misrepresent culture. Prior work has examined cultural bias mainly in text-to-image (T2I) systems, leaving image-to-image (I2I) editors underexplored. We bridge this gap with a unified evaluation across six countries, an 8-category/36-subcategory schema, and era-aware prompts, auditing both T2I generation and I2I editing under a standardized protocol that yields comparable diagnostics. Using open models with fixed settings, we derive cross-country, cross-era, and cross-category evaluations. Our framework combines standard automatic metrics, a culture-aware retrieval-augmented VQA, and expert human judgments collected from native reviewers. To enable reproducibility, we release the complete image corpus, prompts, and configurations. Our study reveals three findings: (1) under country-agnostic prompts, models default to Global-North, modern-leaning depictions that flatten cross-country distinctions; (2) iterative I2I editing erodes cultural fidelity even when conventional metrics remain flat or improve; and (3) I2I models apply superficial cues (palette shifts, generic props) rather than era-consistent, context-aware changes, often retaining source identity for Global-South targets. These results highlight that culture-sensitive edits remain unreliable in current systems. By releasing standardized data, prompts, and human evaluation protocols, we provide a reproducible, culture-centered benchmark for diagnosing and tracking cultural bias in generative image models. Project page: https://seochan99.github.io/ECB

文化偏见图像生成评估框架公平性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。