arXiv:2506.08835cs.CVcs.AI2025-06EMNLP被引 14

首次量化评估文生图模型对文化期待的符合度,发现平均44%的文化线索被忽略。

CulturalFrames: Assessing Cultural Expectation Alignment in Text-to-Image Models and Evaluation Metrics

  • 构建CulturalFrames基准,通过人类标注评估跨文化视觉生成
  • 模型平均遗漏44%文化线索,显性期待遗漏率达68%
  • 现有评测指标与人类判断严重不符,适合关注全球可用性的研究者

文生图模型在视觉内容生成中的普及引发对其能否准确呈现多元文化背景的担忧——文化线索缺失可能导致群体刻板印象并降低实用性。本文首次系统量化文生图模型及评测指标在显性(明确陈述)和隐性(由提示语文化语境暗示)文化期待上的契合度。为此,我们提出CulturalFrames基准,用于严格的人类评估文化表现。该基准涵盖10个国家、5个社会文化领域,包含983个提示词、4个顶尖文生图模型生成的3637张图像,以及超过1万条详细人工标注。结果显示,在不同模型和国家间,文化期待平均被遗漏44%;其中显性期待遗漏率高达68%,隐性期待也平均遗漏49%。此外,我们发现现有文生图评测指标与人类对文化契合度的判断相关性极低,无论其内部推理机制如何。总体而言,研究揭示了关键差距,提供了具体测试平台,并为开发更具文化敏感性的文生图模型与评测方法指明了可行方向。

原文摘要 · Abstract (English)

The increasing ubiquity of text-to-image (T2I) models as tools for visual content generation raises concerns about their ability to accurately represent diverse cultural contexts -- where missed cues can stereotype communities and undermine usability. In this work, we present the first study to systematically quantify the alignment of T2I models and evaluation metrics with respect to both explicit (stated) as well as implicit (unstated, implied by the prompt's cultural context) cultural expectations. To this end, we introduce CulturalFrames, a novel benchmark designed for rigorous human evaluation of cultural representation in visual generations. Spanning 10 countries and 5 socio-cultural domains, CulturalFrames comprises 983 prompts, 3637 corresponding images generated by 4 state-of-the-art T2I models, and over 10k detailed human annotations. We find that across models and countries, cultural expectations are missed an average of 44% of the time. Among these failures, explicit expectations are missed at a surprisingly high average rate of 68%, while implicit expectation failures are also significant, averaging 49%. Furthermore, we show that existing T2I evaluation metrics correlate poorly with human judgments of cultural alignment, irrespective of their internal reasoning. Collectively, our findings expose critical gaps, provide a concrete testbed, and outline actionable directions for developing culturally informed T2I models and metrics that improve global usability.

文生图文化评估人类评估基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。