arXiv:2507.04151cs.CV2025-07

无需人工标注,通过自监督学习实现复杂图像生成的精准控制。

Unlocking Compositional Control: Self-Supervision for LVLM-Based Image Generation

  • 分层自监督训练,让模型自动生成并对齐图像的全局与局部描述。
  • 引入内部组合规划机制,生成细粒度子提示并优化输出一致性。
  • 在多个基准上超越主流模型,适合需要高精度图文一致性的场景。

本文提出层级自监督视觉语言模型(Hi-SSLVLM),一种新型生成模型,显著提升复杂文本到图像合成的效果,尤其在处理组合性挑战性提示时表现突出。传统方法依赖昂贵的人工标注图像-文本对,且难以精确控制细粒度视觉属性与复杂空间关系。Hi-SSLVLM采用双阶段自监督学习策略:第一阶段为多粒度视觉-语言对齐,使大视觉语言模型(LVLM)自主生成并匹配图像的层次化标题(全局与局部),建立无需大量人工标注的深层语义理解;第二阶段为自精炼与引导生成,利用内部组合规划(ICP)机制,由LVLM先制定详细子提示以指导生成过程,并引入新颖的语义一致性损失实现精确输出对齐。在Gemini-2.0-Flash、InternVL3-78B等多维基准上,与Janus-Pro-1B、Stable Diffusion XL 1.0、DeepFloyd IF v1.0、ControlNet-XL等领先基线对比的全面实验表明,Hi-SSLVLM在所有细粒度指标上均表现更优。深入的消融实验验证了各组件的关键作用。人类评估进一步支持定量结果,凸显其在提示忠实度、组合准确性与整体美学质量上的显著提升,标志着开放式文本到图像生成向更高可控性与语义一致性迈进的重要一步。

原文摘要 · Abstract (English)

This paper introduces Hierarchical Self-Supervised LVLM (Hi-SSLVLM), a novel generative model designed to significantly advance text-to-image synthesis, particularly for complex and compositionally challenging prompts. Traditional methods often grapple with the high cost of meticulously curated paired image-text datasets and struggle with precise control over fine-grained visual attributes and intricate spatial relationships. Our Hi-SSLVLM addresses these limitations through a unique two-stage self-supervised learning strategy. The first stage, Multi-Granularity Visual-Language Grounding, enables the Large Vision-Language Model (LVLM) backbone to autonomously generate and align hierarchical captions (global and local) to images, cultivating a deep internal semantic understanding without reliance on extensive human annotation. The second stage, Self-Refinement and Guided Image Generation, leverages this acquired knowledge by an Internal Compositional Planning (ICP) mechanism, where the LVLM first formulates detailed textual sub-prompts to guide the image generation process, complemented by a novel Semantic Consistency Loss for precise output alignment. Comprehensive experiments against leading baselines, including Janus-Pro-1B, Stable Diffusion XL 1.0, DeepFloyd IF v1.0, and ControlNet-XL, on multi-dimensional benchmarks such as Gemini-2.0-Flash and InternVL3-78B, demonstrate Hi-SSLVLM's superior performance across all fine-grained metrics. An in-depth ablation study confirms the critical role of each proposed component. Furthermore, human evaluations corroborate our quantitative findings, highlighting Hi-SSLVLM's enhanced fidelity to prompt, compositional accuracy, and overall aesthetic quality, marking a significant step towards more controllable and semantically consistent open-ended text-to-image generation.

文本生成图像自监督学习可控生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。