让AI生成图像时更准地理解物体属性,减少幻觉。
OSPO: Object-Centric Self-Improving Preference Optimization for Text-to-Image Generation
- 用自生成数据和注意力掩码聚焦物体级对齐
- 在三个基准上显著降低物体幻觉,超越现有方法
- 适合需要精准图文对齐的图像生成研究者
多模态大模型在统一理解与生成方面取得进展,但仍难以实现细粒度的文本-图像对齐,常出现颜色、形状、空间关系等属性错误。以往偏好优化方法如DPO和GRPO计算成本高,虽有自改进方法能自主构建训练数据,但仍忽略物体级语义,导致物体幻觉持续存在。为此,我们提出面向物体中心的自改进偏好优化(OSPO),无需外部数据或模型即可构建物体级偏好数据,并引入基于注意力的物体掩码与物体加权SimPO损失,提升物体特异性保真度。在三个组合式图像生成基准上的实验证明,OSPO显著改善细粒度对齐,减少物体幻觉,性能优于已有自改进方法,甚至超过专门的扩散模型。
原文摘要 · Abstract (English)
Recent advances in Multimodal Large Language Models (MLLMs) have enabled unified multimodal understanding and generation. However, they still struggle with fine-grained text-image alignment, often failing to faithfully depict objects with correct attributes such as color, shape, and spatial relations. To mitigate this issue, previous studies have explored preference optimization methods such as DPO and GRPO, but these approaches incur substantial computational cost, both in constructing preference data and in performing optimization. This has motivated self-improving preference optimization approaches, in which the MLLM autonomously generates its own training data, self-estimates preference feedback, and self-optimizes using the resulting self-constructed preference pairs. However, existing self-improving methods still overlook fine-grained, object-level semantics, allowing object hallucination to persist. To tackle this problem, we propose Object-centric Self-improving Preference Optimization (OSPO), a self-improving framework designed to enhance object-level text-image alignment. OSPO explicitly constructs object-centric preference data without relying on any external data and external models. We also introduce a new approach that leverages attention-based object masks together with an object-weighted SimPO loss to enhance object-specific fidelity. Extensive experiments on three compositional image generation benchmarks demonstrate that OSPO significantly improves fine-grained alignment and reduces object hallucination, outperforming prior self-improving methods and even specialized diffusion-based text-to-image models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。