用结构化推理让模型更懂用户意图,生成图像更准
Re-Align: Structured Reasoning-guided Alignment for In-Context Image Generation and Editing
- 通过分步推理拆解图文指令,避免参考图混淆
- 在相同规模下,生成和编辑效果优于现有方法
- 适合需要精准理解用户指令的图像生成任务
上下文图像生成与编辑(ICGE)允许用户通过交错的图文提示指定视觉概念,要求模型精确理解并忠实执行用户意图。尽管近期统一多模态模型展现出良好的理解能力,但这些优势难以有效迁移至图像生成。本文提出 Re-Align,一种通过结构化推理引导对齐的统一框架,核心是上下文链式思维(IC-CoT),该机制将语义引导与参考关联分离,明确文本目标,缓解参考图像间的混淆问题。此外,Re-Align引入一种有效的强化学习训练方案,利用代理奖励评估结构化推理文本与生成图像之间的对齐度,从而提升模型在 ICGE 任务上的整体表现。大量实验表明,Re-Align 在同等模型规模与资源条件下,显著优于现有对比方法,在图像生成与编辑任务上均取得更好效果。
原文摘要 · Abstract (English)
In-context image generation and editing (ICGE) enables users to specify visual concepts through interleaved image-text prompts, demanding precise understanding and faithful execution of user intent. Although recent unified multimodal models exhibit promising understanding capabilities, these strengths often fail to transfer effectively to image generation. We introduce Re-Align, a unified framework that bridges the gap between understanding and generation through structured reasoning-guided alignment. At its core lies the In-Context Chain-of-Thought (IC-CoT), a structured reasoning paradigm that decouples semantic guidance and reference association, providing clear textual target and mitigating confusion among reference images. Furthermore, Re-Align introduces an effective RL training scheme that leverages a surrogate reward to measure the alignment between structured reasoning text and the generated image, thereby improving the model's overall performance on ICGE tasks. Extensive experiments verify that Re-Align outperforms competitive methods of comparable model scale and resources on both in-context image generation and editing tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。