arXiv:2506.09109cs.CVcs.CL2025-06被引 1

用检索增强评估图像文化相关性,解决生成模型跨文化偏见难测问题。

CAIRe: Cultural Attribution of Images by Retrieval-Augmented Evaluation

  • 基于知识库对图像实体进行文化标注,独立评分每种文化标签。
  • 在罕见文化物品数据集上比基线高22% F1分数。
  • 与人类评分相关性达0.56~0.66,适合评估生成图像文化公平性。

随着文本到图像模型日益普及,确保其在不同文化背景下的公平表现至关重要。现有缓解跨文化偏见的方法常伴随性能下降、事实错误或不当输出等权衡。尽管问题广受关注,但缺乏可靠的测量手段阻碍了进展。为此,我们提出CAIRe,一种评估图像文化相关性的评测指标,可针对用户定义的标签判断文化契合度。该框架将图像中的实体与概念锚定至知识库,利用事实信息对每种文化标签给出独立评分。在使用语言模型构建的手动标注文化显著但稀有物品数据集上,CAIRe相比所有基线提升22% F1分数。此外,我们构建了两个涵盖文化普遍概念的数据集:一个包含文本到图像生成结果,另一个来自自然数据。在5级李克特量表下,CAIRe与人工评分的皮尔逊相关系数分别为0.56和0.66,表明其在多种图像来源中均与人类判断高度一致。

原文摘要 · Abstract (English)

As text-to-image models become increasingly prevalent, ensuring their equitable performance across diverse cultural contexts is critical. Efforts to mitigate cross-cultural biases have been hampered by trade-offs, including a loss in performance, factual inaccuracies, or offensive outputs. Despite widespread recognition of these challenges, an inability to reliably measure these biases has stalled progress. To address this gap, we introduce CAIRe, an evaluation metric that assesses the degree of cultural relevance of an image, given a user-defined set of labels. Our framework grounds entities and concepts in the image to a knowledge base and uses factual information to give independent graded judgments for each culture label. On a manually curated dataset of culturally salient but rare items built using language models, CAIRe surpasses all baselines by 22% F1 points. Additionally, we construct two datasets for culturally universal concepts, one comprising T2I-generated outputs and another retrieved from naturally occurring data. CAIRe achieves Pearson's correlations of 0.56 and 0.66 with human ratings on these sets, based on a 5-point Likert scale of cultural relevance. This demonstrates its strong alignment with human judgment across diverse image sources.

图像评估文化偏见生成模型评测指标

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。