YOLO-Count让图像生成能精准控制物体数量。
YOLO-Count: Differentiable Object Counting for Text-to-Image Generation
- 用新型'基数图'回归目标,适应物体大小和分布差异。
- 在COCO和RefCOCO+数据集上计数准确率超现有方法。
- 适合需要精确数量控制的文本生成图像场景。
我们提出YOLO-Count,一种可微分的开放词汇物体计数模型,既能解决通用计数挑战,又能实现文本到图像(T2I)生成中的精确数量控制。核心贡献是引入‘基数’图这一新型回归目标,可捕捉物体尺寸和空间分布的变化。通过表征对齐与混合强弱监督策略,YOLO-Count弥合了开放词汇计数与T2I生成控制之间的差距。其全可微架构支持基于梯度的优化,可实现准确的物体数量估计,并为生成模型提供细粒度引导。大量实验表明,YOLO-Count在计数精度上达到当前最优水平,同时为T2I系统提供稳健有效的数量控制能力。
原文摘要 · Abstract (English)
We propose YOLO-Count, a differentiable open-vocabulary object counting model that tackles both general counting challenges and enables precise quantity control for text-to-image (T2I) generation. A core contribution is the 'cardinality' map, a novel regression target that accounts for variations in object size and spatial distribution. Leveraging representation alignment and a hybrid strong-weak supervision scheme, YOLO-Count bridges the gap between open-vocabulary counting and T2I generation control. Its fully differentiable architecture facilitates gradient-based optimization, enabling accurate object count estimation and fine-grained guidance for generative models. Extensive experiments demonstrate that YOLO-Count achieves state-of-the-art counting accuracy while providing robust and effective quantity control for T2I systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。