arXiv:2506.15368cs.CVcs.AI2025-06AAAI被引 14

基于文本或图片描述,在视频中自动统计目标物体数量,解决遮挡和重复计数难题。

Open-World Object Counting in Videos

  • 结合图像计数与可提示视频分割跟踪,实现跨帧目标精准识别
  • 在包含遮挡的密集场景中,计数准确率显著优于现有基线模型
  • 适用于动物行为分析、工业质检等需动态物体统计的场景

我们提出视频中开放世界物体计数的新任务:给定文本描述或图像示例以指定目标物体,目标是统计视频中该类物体的所有唯一实例。该任务在存在遮挡和外观相似物体的拥挤场景中尤为挑战,避免重复计数和识别重新出现至关重要。为此,我们提出 CountVid 模型,融合基于图像的计数模型与可提示的视频分割跟踪模型,实现跨视频帧的自动化开放世界物体计数。为评估性能,我们构建了 VideoCount 数据集,源自 TAO 与 MOT20 跟踪数据集,并新增企鹅和金属合金结晶的 X 射线视频。实验表明,CountVid 在该数据集上实现了高精度计数,显著超越强基线模型。VideoCount 数据集、CountVid 模型及全部代码已公开于 https://www.robots.ox.ac.uk/~vgg/research/countvid/。

原文摘要 · Abstract (English)

We introduce a new task of open-world object counting in videos: given a text description, or an image example, that specifies the target object, the objective is to enumerate all the unique instances of the target objects in the video. This task is especially challenging in crowded scenes with occlusions and objects of similar appearance, where avoiding double counting and identifying reappearances is crucial. To this end, we make the following contributions: we introduce a model, CountVid, for this task. It leverages an image-based counting model, and a promptable video segmentation and tracking model, to enable automated open-world object counting across video frames. To evaluate its performance, we introduce VideoCount, a new dataset for this novel task built from the TAO and MOT20 tracking datasets, as well as from videos of penguins and metal alloy crystallization captured by x-rays. Using this dataset, we demonstrate that CountVid provides accurate object counts, and significantly outperforms strong baselines. The VideoCount dataset, the CountVid model, and all the code are available at https://www.robots.ox.ac.uk/~vgg/research/countvid/.

视频计数目标追踪开放世界多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。