arXiv:2501.07086cs.CL2025-01中稿 · ICASSP 2025被引 1

用多语言提示提升大模型图文生成能力,让图像更符合人类偏好。

Boosting Text-To-Image Generation via Multilingual Prompting in Large Multimodal Models

  • 将输入文本翻译成多语言,联合提供原句与译文增强理解。
  • 在3个评测集上表现更优,尤其在细节描述和人类偏好匹配上领先。
  • 适合追求高质量、多样化图像生成的研究者与开发者使用。

以往提升大型多模态模型(LMM)文本到图像(T2I)生成性能的工作主要聚焦于丰富上下文学习(ICL)的输入空间,例如提供少量示例并优化图像描述的详尽性与逻辑性。然而,随着对更复杂灵活图像描述需求的增长,如何在ICL范式下增强模型对输入文本的理解仍是一个关键但未充分探索的问题。本文提出一种新方法PMT2I,通过构建并行多语言提示,利用LMM的多语言能力:将输入文本翻译为多种语言,并同时提供原始文本与译文给模型。在两个LMM和三个基准测试上的实验表明,PMT2I在通用性、组合性及细粒度评估中均表现更优,尤其在人类偏好对齐方面显著领先。此外,由于能生成更多样化的图像,结合重排序方法时,PMT2I显著优于基线提示。代码与多语言数据已开源:https://github.com/takagi97/PMT2I。

原文摘要 · Abstract (English)

Previous work on augmenting large multimodal models (LMMs) for text-to-image (T2I) generation has focused on enriching the input space of in-context learning (ICL). This includes providing a few demonstrations and optimizing image descriptions to be more detailed and logical. However, as demand for more complex and flexible image descriptions grows, enhancing comprehension of input text within the ICL paradigm remains a critical yet underexplored area. In this work, we extend this line of research by constructing parallel multilingual prompts aimed at harnessing the multilingual capabilities of LMMs. More specifically, we translate the input text into several languages and provide the models with both the original text and the translations. Experiments on two LMMs across 3 benchmarks show that our method, PMT2I, achieves superior performance in general, compositional, and fine-grained assessments, especially in human preference alignment. Additionally, with its advantage of generating more diverse images, PMT2I significantly outperforms baseline prompts when incorporated with reranking methods. Our code and parallel multilingual data can be found at https://github.com/takagi97/PMT2I.

文本生成图像多语言提示大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。