arXiv:2603.18001cs.CV2026-03AAAI被引 1

统一生成与理解图像布局,提升准确性和鲁棒性

EchoGen: Cycle-Consistent Learning for Unified Layout-Image Generation and Understanding

  • 通过循环一致性学习联合训练布局生成与图像定位
  • 在多个基准上达到顶尖性能,双任务协同提升效果
  • 适合需要高精度图像生成与语义理解的研究者

本文提出EchoGen,一个统一的布局到图像生成与图像定位框架,能够在生成符合空间关系等文本描述的高质量图像的同时,实现稳健的图像定位。我们认为图像定位具备强大的文本与布局理解能力,可弥补布局生成的不足;而由布局生成的多样化内容又能增强图像定位的鲁棒性。将两任务联合训练于统一模型中,可促进相互提升。然而我们发现该联合训练范式存在优化挑战,导致性能受限。为此,我们提出渐进式训练策略:首先通过并行多任务预训练(PMTP)阶段赋予模型基础能力,利用共享标记加速训练;接着通过双重联合优化(DJO)阶段,利用任务对偶性逐步融合两任务,实现统一优化;最后通过循环强化学习(Cycle RL)阶段,以一致性约束作为奖励,摆脱视觉监督依赖,借助GRPO策略显著提升模型统一能力。大量实验表明,该方法在布局生成与图像定位基准上均达领先水平,并揭示了两任务协同优化带来的显著增益。

原文摘要 · Abstract (English)

In this work, we present EchoGen, a unified framework for layout-to-image generation and image grounding, capable of generating images with accurate layouts and high fidelity to text descriptions (e.g., spatial relationships), while grounding the image robustly at the same time. We believe that image grounding possesses strong text and layout understanding abilities, which can compensate for the corresponding limitations in layout-to-image generation. At the same time, images generated from layouts exhibit high diversity in content, thereby enhancing the robustness of image grounding. Jointly training both tasks within a unified model can promote performance improvements for each. However, we identify that this joint training paradigm encounters several optimization challenges and results in restricted performance. To address these issues, we propose progressive training strategies. First, the Parallel Multi-Task Pre-training (PMTP) stage equips the model with basic abilities for both tasks, leveraging shared tokens to accelerate training. Next, the Dual Joint Optimization (DJO) stage exploits task duality to sequentially integrate the two tasks, enabling unified optimization. Finally, the Cycle RL stage eliminates reliance on visual supervision by using consistency constraints as rewards, significantly enhancing the model's unified capabilities via the GRPO strategy. Extensive experiments demonstrate state-of-the-art results on both layout-to-image generation and image grounding benchmarks, and reveal clear synergistic gains from optimizing the two tasks together.

图像生成布局理解联合训练强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。