arXiv:2502.07455cs.CVcs.AI2025-02NAACL被引 2

构建俄语文化图像生成基准,评估模型对俄罗斯文化元素的准确理解。

RusCode: Russian Cultural Code Benchmark for Text-to-Image Generation

  • 设计19类俄罗斯视觉文化特征,构建含1250条俄英双语提示的数据集。
  • 通过人工评估对比主流生成模型在俄语文化概念上的表现差异。
  • 填补计算机视觉中文化偏见研究空白,适合跨文化生成研究者参考。

文本到图像生成模型在全球用户中日益流行,但许多模型存在强烈的英语文化偏见,忽视或误表其他语言群体、国家和民族的独特性。文化意识缺失会降低生成质量,并引发无意冒犯与偏见传播。与自然语言处理领域相比,计算机视觉中的文化意识研究仍不充分。本文致力于缩小这一差距,提出RusCode基准,用于评估包含俄罗斯文化要素的文本到图像生成质量。我们归纳出19个最能代表俄罗斯视觉文化特征的类别,最终构建包含1250条俄语文本提示及其英文翻译的数据集。这些提示涵盖艺术、流行文化、民间传统、名人姓名、自然物象、科学成就等广泛主题。我们呈现了对主流生成模型在俄罗斯视觉概念表现上的对比人工评估结果。

原文摘要 · Abstract (English)

Text-to-image generation models have gained popularity among users around the world. However, many of these models exhibit a strong bias toward English-speaking cultures, ignoring or misrepresenting the unique characteristics of other language groups, countries, and nationalities. The lack of cultural awareness can reduce the generation quality and lead to undesirable consequences such as unintentional insult, and the spread of prejudice. In contrast to the field of natural language processing, cultural awareness in computer vision has not been explored as extensively. In this paper, we strive to reduce this gap. We propose a RusCode benchmark for evaluating the quality of text-to-image generation containing elements of the Russian cultural code. To do this, we form a list of 19 categories that best represent the features of Russian visual culture. Our final dataset consists of 1250 text prompts in Russian and their translations into English. The prompts cover a wide range of topics, including complex concepts from art, popular culture, folk traditions, famous people's names, natural objects, scientific achievements, etc. We present the results of a human evaluation of the side-by-side comparison of Russian visual concepts representations using popular generative models.

图像生成文化偏见多语言评测基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。