针对布局图重叠生成难题,构建新基准并提出评估指标。
OverLayBench: A Benchmark for Layout-to-Image Generation with Dense Overlaps
- 提出OverLayScore度量重叠复杂度,量化盒子重叠程度。
- 新基准OverLayBench覆盖高重叠场景,打破旧数据集偏差。
- 推出CreatiLayout-AM模型,提升复杂重叠布局生成质量。
尽管布局到图像生成取得进展,但现有方法在存在显著边界框重叠的布局上仍表现不佳。我们识别出两大挑战:(1) 大面积重叠区域,(2) 语义差异极小的重叠实例。通过定性示例和定量分析,我们证明这些因素会降低生成质量。为系统评估该问题,提出OverLayScore,一种量化边界框重叠复杂度的新指标。分析显示,现有基准偏向于OverLayScore值较低的简单情况,限制了其在复杂条件下的评估能力。为此,我们构建OverLayBench,包含高质量标注且在不同OverLayScore水平上分布均衡。作为提升复杂重叠生成性能的初步尝试,我们还提出CreatiLayout-AM,基于精选的非完整掩码数据集微调的模型。我们的工作为真实复杂场景下的鲁棒布局生成奠定了基础。
原文摘要 · Abstract (English)
Despite steady progress in layout-to-image generation, current methods still struggle with layouts containing significant overlap between bounding boxes. We identify two primary challenges: (1) large overlapping regions and (2) overlapping instances with minimal semantic distinction. Through both qualitative examples and quantitative analysis, we demonstrate how these factors degrade generation quality. To systematically assess this issue, we introduce OverLayScore, a novel metric that quantifies the complexity of overlapping bounding boxes. Our analysis reveals that existing benchmarks are biased toward simpler cases with low OverLayScore values, limiting their effectiveness in evaluating model performance under more challenging conditions. To bridge this gap, we present OverLayBench, a new benchmark featuring high-quality annotations and a balanced distribution across different levels of OverLayScore. As an initial step toward improving performance on complex overlaps, we also propose CreatiLayout-AM, a model fine-tuned on a curated amodal mask dataset. Together, our contributions lay the groundwork for more robust layout-to-image generation under realistic and challenging scenarios. Project link: https://mlpc-ucsd.github.io/OverLayBench.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。