arXiv:2608.11002cs.CLcs.AI2026-08中稿 · ACM MM 2026

构建多语言图文生成基准,揭示不同语言生成效果差异

On the Limitations of Cross-Lingual Consistency in Multilingual Text-to-image Generation

论文配图:On the Limitations of Cross-Lingual Consistency in Multilingual Text-to-image Generation
图 1 · 摘自论文原文
  • 构建覆盖10种语言的3.3万条提示的评测基准LingT2I
  • 发现语言间生成质量不均,存在语言依赖性权衡现象
  • 适合关注多语言模型公平性与文化适配的研究者

近年来,文本到图像(T2I)生成取得显著进展,但现有研究主要聚焦于英语场景,跨语言性能差距与语言特异性效应尚未充分探索。为此,我们提出LingT2I基准,涵盖10种常用语言的3.3万条提示,用于评估内容生成与文本渲染中的跨语言影响。基于该基准,我们开展全面的跨语言分析,揭示了语言不平等与评估维度间的语言依赖性权衡。除量化评估外,还发现了多种语言相关的生成模式,表明语言特征及其文化背景系统性影响模型输出。本研究为理解T2I生成中的跨语言行为提供基础,并推动更鲁棒、更具包容性的模型发展。代码与数据集已开源。

原文摘要 · Abstract (English)

Text-to-image (T2I) generation has achieved remarkable progress in recent years. However, existing research has largely focused on English-only settings, leaving cross-lingual performance gaps and language-specific effects insufficiently explored. To fill this gap, we introduce LingT2I, a benchmark covering 10 widely used languages with 33K prompts, designed to evaluate cross-lingual effects in both content generation and text rendering. Building on this benchmark, we conduct a comprehensive cross-lingual analysis, uncovering linguistic inequality and language-dependent trade-offs across evaluation dimensions. Beyond quantitative evaluation, we further reveal a range of language-dependent generation patterns, highlighting how linguistic factors and their corresponding cultural contexts systematically impact model outputs. Our benchmark and analysis provide a foundation for studying cross-lingual behavior in T2I generation and facilitate the development of more robust and inclusive models. Code and dataset are available at https://github.com/RISys-Lab/LingT2I.

多语言图文生成基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。