无需训练即可精准生成指定数量的物体,特别适合高密度场景。
CountLoop: Training-Free High-Instance Image Generation via Iterative Agent Guidance
- 通过迭代式规划与评价,用视觉语言模型控制生成布局。
- 在高密度场景中计数误差降低57%,空间布局质量最优。
- 适合需要精确物体数量和位置控制的图像生成任务。
扩散模型在生成逼真图像方面表现优异,但在高密度场景中难以精确控制物体数量。我们提出COUNTLOOP,一种无需训练的框架,通过迭代的结构化反馈实现精确实例控制。该方法交替进行生成与评估:基于视觉语言模型的规划器生成结构化场景布局,而基于视觉语言模型的评判器则提供关于物体数量、空间排列和视觉质量的明确反馈,以迭代优化布局。实例驱动的注意力掩码和累积注意力组合进一步防止语义泄露,确保在高度重叠场景中仍能保持清晰的物体区分。在COCO-Count、T2I-CompBench以及两个新提出的高实例基准测试上,COUNTLOOP将计数误差最多降低57%,并在所有基准上达到最高或相当的空间质量得分,同时保持图像的逼真性。
原文摘要 · Abstract (English)
Diffusion models excel at photorealistic synthesis but struggle with precise object counts, especially in high-density settings. We introduce COUNTLOOP, a training-free framework that achieves precise instance control through iterative, structured feedback. Our method alternates between synthesis and evaluation: a VLM-based planner generates structured scene layouts, while a VLM-based critic provides explicit feedback on object counts, spatial arrangements, and visual quality to refine the layout iteratively. Instance-driven attention masking and cumulative attention composition further prevent semantic leakage, ensuring clear object separation even in densely occluded scenes. Evaluations on COCO-Count, T2I-CompBench, and two newly introduced high instance benchmarks show that COUNTLOOP reduces counting error by up to 57% and achieves the highest or comparable spatial quality scores across all benchmarks, while maintaining photorealism.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。