arXiv:2502.16902cs.CVcs.AI2025-02NAACL被引 13

让AI更懂文化细节,精准生成非西方文化物品图像

Culture-TRIP: Culturally-Aware Text-to-Image Generation with Iterative Prompt Refinement

  • 通过检索维基百科和网络信息,迭代优化包含文化名词的提示词
  • 用户调研显示,对韩式餐具等冷门文化物品生成准确率提升明显
  • 适合关注跨文化内容生成的研究者与创作者使用

文本到图像模型(如Stable Diffusion)虽在语义对齐方面显著进步,但仍难以准确生成非西方文化中不常见或代表性不足的概念或物体图像,例如韩式餐具‘hangari’。本文提出一种名为Culture-TRIP的新方法,通过迭代提示词精炼实现文化感知的文本到图像生成。该方法首先检索文化名词相关的文化背景与视觉细节,再基于文化标准与大语言模型进行多轮提示词优化与评估。优化过程利用维基百科及网络信息源。针对来自八个国家的66名参与者开展的用户调研表明,该方法显著提升了图像与提示词之间的对齐度,尤其在冷门文化名词的生成上表现更优。

原文摘要 · Abstract (English)

Text-to-Image models, including Stable Diffusion, have significantly improved in generating images that are highly semantically aligned with the given prompts. However, existing models may fail to produce appropriate images for the cultural concepts or objects that are not well known or underrepresented in western cultures, such as `hangari' (Korean utensil). In this paper, we propose a novel approach, Culturally-Aware Text-to-Image Generation with Iterative Prompt Refinement (Culture-TRIP), which refines the prompt in order to improve the alignment of the image with such culture nouns in text-to-image models. Our approach (1) retrieves cultural contexts and visual details related to the culture nouns in the prompt and (2) iteratively refines and evaluates the prompt based on a set of cultural criteria and large language models. The refinement process utilizes the information retrieved from Wikipedia and the Web. Our user survey, conducted with 66 participants from eight different countries demonstrates that our proposed approach enhances the alignment between the images and the prompts. In particular, C-TRIP demonstrates improved alignment between the generated images and underrepresented culture nouns. Resource can be found at https://shane3606.github.io/Culture-TRIP.

文本生成文化感知提示优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。