arXiv:2603.24965cs.CVcs.AI2026-03被引 1

用可解释的奖励机制让图像生成自己纠错,提升细节和空间关系准确性。

Self-Corrected Image Generation with Explainable Latent Rewards

  • 通过多模态大模型生成反馈,指导图像在潜在空间自我修正。
  • 在多个任务中显著提升语义对齐度与视觉保真度,且不破坏原有生成风格。
  • 适合需要高精度控制的图像生成与编辑场景,如设计、艺术创作。

尽管文本到图像生成已取得显著进展,但对复杂提示词(尤其是细粒度语义与空间关系)的对齐仍具挑战性,根源在于生成过程的前馈特性——模型需在未完全理解输出的情况下预判对齐效果。相比之下,图像评估更为可行。基于这一不对称性,我们提出 xLARD:一种利用多模态大语言模型生成可解释潜在奖励的自纠正框架。xLARD 引入轻量级校正器,依据模型生成的参考样本结构化反馈,优化潜在表示。核心是将潜在编辑映射为可解释的奖励信号,实现从非可微图像评估到连续潜在层指导的可微映射。该机制使模型能在生成过程中实现理解、评估与自我修正。跨多种生成与编辑任务的实验表明,xLARD 显著提升了语义对齐度与视觉保真度,同时保持原始生成先验。代码已开源:https://yinyiluo.github.io/xLARD/。

原文摘要 · Abstract (English)

Despite significant progress in text-to-image generation, aligning outputs with complex prompts remains challenging, particularly for fine-grained semantics and spatial relations. This difficulty stems from the feed-forward nature of generation, which requires anticipating alignment without fully understanding the output. In contrast, evaluating generated images is more tractable. Motivated by this asymmetry, we propose xLARD, a self-correcting framework that uses multimodal large language models to guide generation through Explainable LAtent RewarDs. xLARD introduces a lightweight corrector that refines latent representations based on structured feedback from model-generated references. A key component is a differentiable mapping from latent edits to interpretable reward signals, enabling continuous latent-level guidance from non-differentiable image-level evaluations. This mechanism allows the model to understand, assess, and correct itself during generation. Experiments across diverse generation and editing tasks show that xLARD improves semantic alignment and visual fidelity while maintaining generative priors. Code is available at https://yinyiluo.github.io/xLARD/.

图像生成自纠正可解释性多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。