arXiv:2508.04732cs.LGcs.GR2025-08被引 1

用视觉语言模型迭代优化文本生成图像,提升细节控制与语义一致性。

LumiGen: An LVLM-Enhanced Iterative Framework for Fine-Grained Text-to-Image Generation

  • 引入视觉语言模型做闭环反馈,动态修正生成图像
  • 在LongBench-T2I上平均得分3.08,优于现有方法
  • 特别改善文字渲染与姿态表达,适合精细图像生成任务

文本到图像(T2I)生成虽因扩散模型取得显著进展,但在处理复杂指令、实现细粒度内容控制及保持深层语义一致性方面仍面临挑战。现有T2I模型在文字精准呈现、姿态准确生成和复杂组合一致性等任务上表现不足。与此同时,视觉语言模型(LVLM)在跨模态理解与指令遵循方面展现出强大能力。本文提出LumiGen,一种基于LVLM增强的迭代框架,通过闭环的LVLM驱动反馈机制,显著提升T2I模型在细粒度控制场景下的表现。LumiGen包含智能提示解析与增强(IPPA)模块,用于主动优化提示词;以及迭代视觉反馈与优化(IVFR)模块,充当‘视觉批评者’,持续修正并优化生成图像。在具有挑战性的LongBench-T2I基准上,LumiGen取得3.08的平均分,超越当前最优基线。尤其在文字渲染与姿态表达等关键维度上表现显著提升,验证了LVLM融合对更可控、更高品质图像生成的有效性。

原文摘要 · Abstract (English)

Text-to-Image (T2I) generation has made significant advancements with diffusion models, yet challenges persist in handling complex instructions, ensuring fine-grained content control, and maintaining deep semantic consistency. Existing T2I models often struggle with tasks like accurate text rendering, precise pose generation, or intricate compositional coherence. Concurrently, Vision-Language Models (LVLMs) have demonstrated powerful capabilities in cross-modal understanding and instruction following. We propose LumiGen, a novel LVLM-enhanced iterative framework designed to elevate T2I model performance, particularly in areas requiring fine-grained control, through a closed-loop, LVLM-driven feedback mechanism. LumiGen comprises an Intelligent Prompt Parsing & Augmentation (IPPA) module for proactive prompt enhancement and an Iterative Visual Feedback & Refinement (IVFR) module, which acts as a "visual critic" to iteratively correct and optimize generated images. Evaluated on the challenging LongBench-T2I Benchmark, LumiGen achieves a superior average score of 3.08, outperforming state-of-the-art baselines. Notably, our framework demonstrates significant improvements in critical dimensions such as text rendering and pose expression, validating the effectiveness of LVLM integration for more controllable and higher-quality image generation.

文本生成图像视觉语言模型迭代优化细粒度控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。