arXiv:2508.03199cs.CL2025-08EMNLP被引 4

语言语法性别会显著影响图文模型的视觉生成结果。

Beyond Content: How Grammatical Gender Shapes Visual Representation in Text-to-Image Models

  • 通过跨语言对比实验,研究语法性别如何塑造图像生成。
  • 语法男性标记使男性形象出现率升至73%,女性标记达38%。
  • 揭示了语言结构本身是影响多语言生成公平性的新因素。

图文生成模型中的偏见研究主要聚焦于人口统计特征和刻板印象,却忽略了语言语法性别对视觉表征的影响。本文构建了一个跨语言基准,考察语法性别与刻板性别认知不一致的词汇(如法语中语法阴性但指代‘守卫’这一典型男性概念的‘une sentinelle’)。数据集涵盖五种有语法性别语言(法语、西班牙语、德语、意大利语、俄语)和两种无性别控制语言(英语、中文),共800个独特提示,生成28,800张图像,覆盖三种主流T2I模型。分析显示,语法性别显著影响生成结果:语法阳性标记使男性形象平均占比达73%(英语为22%),语法阴性标记使女性形象占38%(英语为28%)。该效应随语言资源丰富度和模型架构系统性变化,高资源语言中表现更明显。研究证明,语言结构本身而非仅内容,塑造了多语言多模态系统的生成输出,为理解偏见与公平性提供了新维度。

原文摘要 · Abstract (English)

Research on bias in Text-to-Image (T2I) models has primarily focused on demographic representation and stereotypical attributes, overlooking a fundamental question: how does grammatical gender influence visual representation across languages? We introduce a cross-linguistic benchmark examining words where grammatical gender contradicts stereotypical gender associations (e.g., ``une sentinelle'' - grammatically feminine in French but referring to the stereotypically masculine concept ``guard''). Our dataset spans five gendered languages (French, Spanish, German, Italian, Russian) and two gender-neutral control languages (English, Chinese), comprising 800 unique prompts that generated 28,800 images across three state-of-the-art T2I models. Our analysis reveals that grammatical gender dramatically influences image generation: masculine grammatical markers increase male representation to 73% on average (compared to 22% with gender-neutral English), while feminine grammatical markers increase female representation to 38% (compared to 28% in English). These effects vary systematically by language resource availability and model architecture, with high-resource languages showing stronger effects. Our findings establish that language structure itself, not just content, shapes AI-generated visual outputs, introducing a new dimension for understanding bias and fairness in multilingual, multimodal systems.

文本生成图像语言偏见语法性别多语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。