arXiv:2412.05818cs.CVcs.AI2024-12CVPR被引 19

让AI模型自我纠错,提升图文生成准确率

SILMM: Self-Improving Large Multimodal Models for Compositional Text-to-Image Generation

  • 通过自反馈机制让模型迭代优化图文对齐
  • 在多个基准测试中性能提升超30%
  • 适用于各种视觉表示的模型,无需人工标注

大型多模态模型(LMMs)在多模态理解与生成方面表现出色,推动了文本到图像生成的发展。然而,在组合场景下实现精准的文本-图像对齐仍具挑战。现有方法如分步布局规划或基于人类/人工智能反馈的学习,高度依赖提示工程、昂贵的人工标注和持续升级,限制了灵活性与可扩展性。本文提出一种模型无关的迭代自改进框架SILMM,使LMM能够生成有益且可扩展的自反馈,并通过直接偏好优化(DPO)优化文本-图像对齐。DPO适用于以离散视觉标记为中间表示的LMM;而对于具有连续视觉特征的LMM,由于生成概率难以获取,适用性受限。为此,我们提出一种多样性机制以获得多样化表示,并设计基于核函数的连续DPO进行对齐。在三个组合式文本到图像生成基准上的大量实验验证了SILMM的有效性与优越性,在T2I-CompBench++上性能提升超过30%,在DPG-Bench上提升约20%。

原文摘要 · Abstract (English)

Large Multimodal Models (LMMs) have demonstrated impressive capabilities in multimodal understanding and generation, pushing forward advancements in text-to-image generation. However, achieving accurate text-image alignment for LMMs, particularly in compositional scenarios, remains challenging. Existing approaches, such as layout planning for multi-step generation and learning from human feedback or AI feedback, depend heavily on prompt engineering, costly human annotations, and continual upgrading, limiting flexibility and scalability. In this work, we introduce a model-agnostic iterative self-improvement framework (SILMM) that can enable LMMs to provide helpful and scalable self-feedback and optimize text-image alignment via Direct Preference Optimization (DPO). DPO can readily applied to LMMs that use discrete visual tokens as intermediate image representations; while it is less suitable for LMMs with continuous visual features, as obtaining generation probabilities is challenging. To adapt SILMM to LMMs with continuous features, we propose a diversity mechanism to obtain diverse representations and a kernel-based continuous DPO for alignment. Extensive experiments on three compositional text-to-image generation benchmarks validate the effectiveness and superiority of SILMM, showing improvements exceeding 30% on T2I-CompBench++ and around 20% on DPG-Bench.

图文生成自进化多模态优化算法

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。