arXiv:2411.19534cs.CVcs.LG2024-11中稿 · WACV 2026被引 3

无需重训练,让文生图模型跨域精准数物体。

QUOTA: Quantifying Objects with Text-to-Image Models for Any Domain

  • 用双循环元学习优化通用提示词,实现跨域无感适配。
  • 在未见领域中准确计数,且对训练外物体类仍保持精度。
  • 适合需要快速部署到新场景的生成式图像应用开发者。

我们解决通过生成式文生图模型进行物体数量统计的问题。不同于为每个新图像领域重新训练模型(导致计算成本高、扩展性差),我们首次从领域无关视角出发,提出QUOTA——一种无需重训练即可在未见领域实现有效物体计数的优化框架。该方法采用双循环元学习策略优化领域不变提示词,并融合提示学习与可学习计数、领域令牌机制,以捕捉风格差异并维持准确性,即使面对训练中未出现的物体类别亦然。我们构建了一个专用于评估文本到图像生成中领域泛化能力的物体计数新基准,支持对跨域适应性和计数精度的严格测试。大量实验表明,QUOTA在物体计数准确率与语义一致性方面均优于传统模型,为任意领域的高效可扩展文生图生成树立了新标准。

原文摘要 · Abstract (English)

We tackle the problem of quantifying the number of objects by a generative text-to-image model. Rather than retraining such a model for each new image domain of interest, which leads to high computational costs and limited scalability, we are the first to consider this problem from a domain-agnostic perspective. We propose QUOTA, an optimization framework for text-to-image models that enables effective object quantification across unseen domains without retraining. It leverages a dual-loop meta-learning strategy to optimize a domain-invariant prompt. Further, by integrating prompt learning with learnable counting and domain tokens, our method captures stylistic variations and maintains accuracy, even for object classes not encountered during training. For evaluation, we adopt a new benchmark specifically designed for object quantification in domain generalization, enabling rigorous assessment of object quantification accuracy and adaptability across unseen domains in text-to-image generation. Extensive experiments demonstrate that QUOTA outperforms conventional models in both object quantification accuracy and semantic consistency, setting a new benchmark for efficient and scalable text-to-image generation for any domain.

文生图物体计数跨域泛化提示优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。