用偏好优化解决前景图像生成中的空间错乱问题。
InpaintDPO: Mitigating Spatial Relationship Hallucinations in Foreground-conditioned Inpainting via Diverse Preference Optimization
- 基于直接偏好优化,聚焦背景空间合理性建模。
- 提出掩码策略避免前景干扰,提升背景关系真实性。
- 适合关注图像生成逻辑一致性与边界融合的研究者。
前景条件图像修复旨在根据给定前景主体和文本提示生成协调的背景,是可控图像生成的重要方向。当前方法常出现前景与背景之间的空间关系幻觉,如尺度失真、位置错误和视角不合理。由于空间合理性具有主观性,难以量化,传统基于奖励的RLHF方法难以适用。为此,本文提出InpaintDPO,首个面向前景条件修复中空间合理性的直接偏好优化框架,确保前景与背景元素间合理的空间关系。为解决标准DPO中相同前景在胜负样本对下引发的梯度冲突,提出MaskDPO,仅对背景区域进行偏好优化,同时保留前景区域的修复损失以保障前景完整性。为增强前景-背景交界处的一致性,提出条件非对称偏好优化,通过差异化裁剪采样配对并实施全局偏好优化,提升上下文感知能力与边界连贯性。最后,基于优质胜出样本共享合理空间关系的观察,提出共享共性偏好优化,强化模型对高质量样本中共性空间规律的理解,进一步促进空间合理性统一。
原文摘要 · Abstract (English)
Foreground-conditioned inpainting, which aims at generating a harmonious background for a given foreground subject based on the text prompt, is an important subfield in controllable image generation. A common challenge in current methods, however, is the occurrence of Spatial Relationship Hallucinations between the foreground subject and the generated background, including inappropriate scale, positional relationships, and viewpoints. Critically, the subjective nature of spatial rationality makes it challenging to quantify, hindering the use of traditional reward-based RLHF methods. To address this issue, we propose InpaintDPO, the first Direct Preference Optimization (DPO) based framework dedicated to spatial rationality in foreground-conditioned inpainting, ensuring plausible spatial relationships between foreground and background elements. To resolve the gradient conflicts in standard DPO caused by identical foreground in win-lose pairs, we propose MaskDPO, which confines preference optimization exclusively to the background to enhance background spatial relationships, while retaining the inpainting loss in the foreground region for robust foreground preservation. To enhance coherence at the foreground-background boundary, we propose Conditional Asymmetric Preference Optimization, which samples pairs with differentiated cropping operations and applies global preference optimization to promote contextual awareness and enhance boundary coherence. Finally, based on the observation that winning samples share a commonality in plausible spatial relationships, we propose Shared Commonality Preference Optimization to enhance the model's understanding of spatial commonality across high-quality winning samples, further promoting shared spatial rationality.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。