构建首个跨文化图文生成基准,揭示主流模型在多元文化场景下的表现差异。
When Cultures Meet: Multicultural Text-to-Image Generation
- 提出多代理框架MosAIG,用不同文化角色的LLM协作生成更精准的文化场景
- 在包含9000张图像的多国数据集上验证,语言与人群差异导致生成质量不均
- 适合关注跨文化生成、公平性与多模态对齐的研究者和开发者
文本到图像生成模型在同质文化场景中表现优异,但在融合不同文化背景的人物与地标等多元文化场景中的能力仍缺乏研究。本文提出跨文化图文生成新任务,并构建首个相关基准数据集。该数据集涵盖5个国家、3个年龄组、2种性别、25个历史地标及5种语言,共9000张图像。基于此,我们从对齐性、图像质量、美学、知识掌握与公平性等多个维度分析了当前领先模型的表现。作为一种文化信息组合策略,本文探索了多代理框架MosAIG,利用具有不同文化身份的LLM提升跨文化图像生成效果。结果表明,更丰富的提示设计可显著提升图像质量与文化准确性,但各语言及人口群体间仍存在明显差距。代码与数据已开源至https://github.com/AIM-SCU/MosAIG。
原文摘要 · Abstract (English)
Text-to-image generation models have achieved strong performance in culturally homogeneous settings, yet their ability to generate multicultural scenes, where people and landmarks originate from different cultures, remains largely unexplored. We introduce multicultural text-to-image generation as a new task and present the first benchmark designed to study this setting. Our dataset contains 9,000 images spanning five countries, three age groups, two genders, 25 historical landmarks, and five languages. Using this benchmark, we analyze the behavior of state-of-the-art text-to-image models across multiple dimensions, including alignment, image quality, aesthetics, knowledge, and fairness. As one strategy for composing cultural and demographic information, we explore MosAIG, a Multi-Agent framework that enhances multicultural Image Generation by leveraging LLMs with distinct cultural personas. Our analysis shows that richer prompt composition can improve image quality and cultural grounding compared to simple prompts, while revealing substantial disparities across languages and demographic groups. We release our dataset and code at https://github.com/AIM-SCU/MosAIG.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。