图像生成模型会错误地按文本顺序排列物体,导致布局出错。
Order Is Not Layout: Order-to-Space Bias in Image Generation
- 通过对比同义但顺序不同的提示词,发现模型受文本顺序误导。
- 多数主流模型存在此偏差,且早期布局阶段就已形成。
- 针对性微调或早期干预可有效降低偏差,保持生成质量。
我们研究现代图像生成模型中一种系统性偏差:文本中实体的提及顺序会虚假地决定空间布局和实体-角色绑定。我们称此现象为顺序到空间偏差(Order-to-Space Bias, OTS),并证明它在文生图和图生图任务中普遍存在,常覆盖真实语义线索,导致布局错误或角色错配。为量化该偏差,我们提出OTS-Bench,通过仅改变实体顺序的成对提示词,评估模型在一致性与正确性两个维度的表现。实验表明,该偏差广泛存在于现代图像生成模型中,且主要由数据驱动,在布局形成的早期阶段即已显现。基于此发现,我们证明针对性微调和早期干预策略可显著降低偏差,同时保持生成质量。
原文摘要 · Abstract (English)
We study a systematic bias in modern image generation models: the mention order of entities in text spuriously determines spatial layout and entity--role binding. We term this phenomenon Order-to-Space Bias (OTS) and show that it arises in both text-to-image and image-to-image generation, often overriding grounded cues and causing incorrect layouts or swapped assignments. To quantify OTS, we introduce OTS-Bench, which isolates order effects with paired prompts differing only in entity order and evaluates models along two dimensions: homogenization and correctness. Experiments show that Order-to-Space Bias (OTS) is widespread in modern image generation models, and provide evidence that it is primarily data-driven and manifests during the early stages of layout formation. Motivated by this insight, we show that both targeted fine-tuning and early-stage intervention strategies can substantially reduce OTS, while preserving generation quality.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。