arXiv:2506.05501cs.CV2025-06被引 15

通过强化学习提升文本图像生成中的细粒度对齐精度

FocusDiff: Advancing Fine-Grained Text-Image Alignment for Autoregressive Visual Generation through RL

  • 用强化学习聚焦相似文本间的细微语义差异
  • 在新基准PairComp上显著优于现有方法
  • 适合需要精准控制图像细节的研究者

近期研究将自回归范式拓展至文本到图像生成,性能已接近扩散模型。然而,我们提出的PairComp基准测试显示,现有模型在细粒度语义对齐方面表现不佳,难以精确控制视觉标记。为此,我们提出FocusDiff,通过关注相似文本-图像对之间的细微差异来增强细粒度对齐。我们构建了一个包含整体表达相似但局部语义不同的配对文本与图像的新数据集,并引入一种新颖的强化学习算法,以强调这些细粒度语义差异,实现期望的图像生成。该方法在现有文本到图像基准上达到最先进性能,在PairComp上显著超越先前方法。

原文摘要 · Abstract (English)

Recent studies extend the autoregression paradigm to text-to-image generation, achieving performance comparable to diffusion models. However, our new PairComp benchmark -- featuring test cases of paired prompts with similar syntax but different fine-grained semantics -- reveals that existing models struggle with fine-grained text-image alignment thus failing to realize precise control over visual tokens. To address this, we propose FocusDiff, which enhances fine-grained text-image semantic alignment by focusing on subtle differences between similar text-image pairs. We construct a new dataset of paired texts and images with similar overall expressions but distinct local semantics, further introducing a novel reinforcement learning algorithm to emphasize such fine-grained semantic differences for desired image generation. Our approach achieves state-of-the-art performance on existing text-to-image benchmarks and significantly outperforms prior methods on PairComp.

自回归生成细粒度对齐强化学习文本图像生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。