arXiv:2608.28206cs.CVcs.DB2026-08

提出新基准与评估方法,揭示图像生成中计数失败的根源。

NumBench: Diagnosing Counting Failures in Text-to-Image Models

论文配图:NumBench: Diagnosing Counting Failures in Text-to-Image Models
图 1 · 摘自论文原文
  • 构建64万条带控变量的提示数据集,系统测试计数能力
  • 发现超过50个物体时性能骤降,布局影响显著
  • 适合研究生成模型计数缺陷的学者与工程师

文本到图像(T2I)模型常生成错误的物体数量,但现有基准规模小且控制弱。我们提出 extbf{NumBench},一个包含64万条提示、覆盖1600类、计数范围1至100的基准。其因子设计同时变化物体构成、空间引导和外观条件,并平衡计数与类别暴露。我们还构建了一个过程模型,其中请求的实例竞争有限可解析图像区域。该模型预测在低占用率下存在近二次方的碰撞缺陷,并显示协调布局可减少此问题。为实现可扩展评估,我们提出置信加权数值精度分数( extit{CWNPS}),整合三个校准检测器并剔除不确定预测。在五个商用系统、两个开源模型及两种专用计数方法上,性能随请求数量急剧下降;所有方法在50以上物体时均表现薄弱。计数范围影响最大,其次为布局与构图。网格引导效果最强,与协调性预测一致,但未确立碰撞为唯一原因。1.44万张图像的人类研究支持自动化评估在计数50以内的有效性,243条自然语言提示的结果表明方法具备跨模板迁移能力。

原文摘要 · Abstract (English)

Text-to-image (T2I) models often generate the wrong number of objects, yet existing benchmarks are too small or weakly controlled to explain why. We introduce \textbf{NumBench}, a benchmark of 640{,}000 prompts spanning 1{,}600 categories and counts from 1 to 100. Its factorial design varies object composition, spatial guidance, and appearance conditions while balancing counts and category exposure. We also develop a process model in which requested instances compete for a finite set of resolvable image regions. The model predicts a near-quadratic collision deficit at low occupancy and shows how coordinated placement reduces it. For scalable evaluation, we propose the Confidence-Weighted Numeric Precision Score (\cwnps), which aggregates three calibrated detectors and discounts uncertain proposals. Across five commercial systems, two open models, and two specialized counting methods, performance declines sharply with requested count; all evaluated methods are weak above 50 objects. Count range has the largest measured effect, followed by layout and composition. Grid guidance is strongest among guided layouts, consistent with the coordination prediction, although the analysis does not establish collision as the sole cause. A 14{,}400-image human study supports automated evaluation through count 50, while results on 243 natural-language prompts show transfer beyond NumBench templates.

图像生成计数误差评估基准模型诊断

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。