arXiv:2503.06884cs.CVcs.AI2025-03被引 20

文生图扩散模型无法准确计数,优化提示词也无济于事。

Text-to-Image Diffusion Models Cannot Count, and Prompt Refinement Cannot Help

论文配图:Text-to-Image Diffusion Models Cannot Count, and Prompt Refinement Cannot Help
图 1 · 摘自论文原文
  • 构建新基准T2ICountBench,精准评估模型计数能力。
  • 所有主流模型计数准确率随物体数量增加急剧下降。
  • 提示词优化对提升计数精度无效,暴露模型本质缺陷。

生成建模是当前人工智能领域最重要的问题之一,文本到图像生成已产生广泛现实影响。扩散模型在该任务中取得显著成功,成为主流方案。然而,这些模型在遵循用户指令中的数量约束方面存在根本性局限,常生成错误数量的物体。尽管已有研究提及此问题,但缺乏系统、严谨的评估。为此,我们提出T2ICountBench,一个新型基准,用于严格评估最先进的文生图扩散模型的计数能力。该基准涵盖多种开源与私有模型,明确分离计数性能与其他能力,提供结构化难度层级,并结合人工评估确保可靠性。大量实验表明,所有先进模型均无法正确生成指定数量的物体,且准确率随物体数量增加而显著下降。此外,对提示词优化的探索发现,此类简单干预通常无法改善计数准确性。研究揭示了扩散模型在数值理解上的内在挑战,并指明未来改进方向。

原文摘要 · Abstract (English)

Generative modeling is widely regarded as one of the most essential problems in today's AI community, with text-to-image generation having gained unprecedented real-world impacts. Among various approaches, diffusion models have achieved remarkable success and have become the de facto solution for text-to-image generation. However, despite their impressive performance, these models exhibit fundamental limitations in adhering to numerical constraints in user instructions, frequently generating images with an incorrect number of objects. While several prior works have mentioned this issue, a comprehensive and rigorous evaluation of this limitation remains lacking. To address this gap, we introduce T2ICountBench, a novel benchmark designed to rigorously evaluate the counting ability of state-of-the-art text-to-image diffusion models. Our benchmark encompasses a diverse set of generative models, including both open-source and private systems. It explicitly isolates counting performance from other capabilities, provides structured difficulty levels, and incorporates human evaluations to ensure high reliability. Extensive evaluations with T2ICountBench reveal that all state-of-the-art diffusion models fail to generate the correct number of objects, with accuracy dropping significantly as the number of objects increases. Additionally, an exploratory study on prompt refinement demonstrates that such simple interventions generally do not improve counting accuracy. Our findings highlight the inherent challenges in numerical understanding within diffusion models and point to promising directions for future improvements.

文生图扩散模型计数能力基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。