arXiv:2512.03233cs.CV2025-12被引 1

用GPT-4o和GPT-5仅靠文本提示实现零样本物体计数,效果媲美顶尖方法。

Object Counting with GPT-4o and GPT-5: A Comparative Study

  • 仅用文本提示,不依赖标注数据或视觉样例,实现零样本计数。
  • 在FSC-147上性能接近甚至超过当前最先进零样本方法。
  • 验证多模态大模型具备无需监督的物体计数推理能力,适合探索新场景。

零样本物体计数旨在估计视觉模型在训练中从未见过的新类别物体实例数量。现有方法通常需要大量标注数据,并常依赖视觉样例引导计数过程。然而,大语言模型(LLMs)具备强大的推理与数据理解能力,暗示其可用于无需任何监督的计数任务。本文旨在利用两种多模态大模型GPT-4o与GPT-5,仅通过文本提示实现零样本物体计数。我们在FSC-147与CARPK数据集上评估两模型并进行对比分析。结果表明,两模型在FSC-147上的表现可与当前最先进的零样本方法相媲美,某些情况下甚至更优。

原文摘要 · Abstract (English)

Zero-shot object counting attempts to estimate the number of object instances belonging to novel categories that the vision model performing the counting has never encountered during training. Existing methods typically require large amount of annotated data and often require visual exemplars to guide the counting process. However, large language models (LLMs) are powerful tools with remarkable reasoning and data understanding abilities, which suggest the possibility of utilizing them for counting tasks without any supervision. In this work we aim to leverage the visual capabilities of two multi-modal LLMs, GPT-4o and GPT-5, to perform object counting in a zero-shot manner using only textual prompts. We evaluate both models on the FSC-147 and CARPK datasets and provide a comparative analysis. Our findings show that the models achieve performance comparable to the state-of-the-art zero-shot approaches on FSC-147, in some cases, even surpass them.

零样本计数多模态大模型GPT-4oGPT-5

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。