arXiv:2508.16644cs.CV2025-08被引 3

无需训练即可精准生成指定数量的物体,特别适合高密度场景。

CountLoop: Training-Free High-Instance Image Generation via Iterative Agent Guidance

  • 通过迭代式规划与评价,用视觉语言模型控制生成布局。
  • 在高密度场景中计数误差降低57%,空间布局质量最优。
  • 适合需要精确物体数量和位置控制的图像生成任务。

扩散模型在生成逼真图像方面表现优异,但在高密度场景中难以精确控制物体数量。我们提出COUNTLOOP,一种无需训练的框架,通过迭代的结构化反馈实现精确实例控制。该方法交替进行生成与评估:基于视觉语言模型的规划器生成结构化场景布局,而基于视觉语言模型的评判器则提供关于物体数量、空间排列和视觉质量的明确反馈,以迭代优化布局。实例驱动的注意力掩码和累积注意力组合进一步防止语义泄露,确保在高度重叠场景中仍能保持清晰的物体区分。在COCO-Count、T2I-CompBench以及两个新提出的高实例基准测试上,COUNTLOOP将计数误差最多降低57%,并在所有基准上达到最高或相当的空间质量得分,同时保持图像的逼真性。

原文摘要 · Abstract (English)

Diffusion models excel at photorealistic synthesis but struggle with precise object counts, especially in high-density settings. We introduce COUNTLOOP, a training-free framework that achieves precise instance control through iterative, structured feedback. Our method alternates between synthesis and evaluation: a VLM-based planner generates structured scene layouts, while a VLM-based critic provides explicit feedback on object counts, spatial arrangements, and visual quality to refine the layout iteratively. Instance-driven attention masking and cumulative attention composition further prevent semantic leakage, ensuring clear object separation even in densely occluded scenes. Evaluations on COCO-Count, T2I-CompBench, and two newly introduced high instance benchmarks show that COUNTLOOP reduces counting error by up to 57% and achieves the highest or comparable spatial quality scores across all benchmarks, while maintaining photorealism.

图像生成扩散模型实例控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。