arXiv:2506.17186cs.CV2025-06

轻量级多目标追踪器,支持单目与双目摄像头

YASMOT: Yet another stereo image multi-object tracker

  • 基于深度学习检测器输出,实现跨时序目标追踪
  • 支持单目/双目输入,可生成检测器集成共识结果
  • 适合需要持续追踪与身份保持的视觉任务

目前已有多种基于深度学习的流行目标检测器,可分析图像并提取物体的位置和类别标签。对于图像时序序列(如视频或一系列静态图像),在时间上追踪物体并保持其身份,有助于提升检测性能,并对许多下游任务(如行为分类与预测、总量估计)至关重要。本文提出 yasmot,一种轻量且灵活的目标追踪器,可处理主流目标检测器的输出,从单目或双目相机配置中追踪物体。此外,它还具备从多个检测器集合中生成共识检测的功能。

原文摘要 · Abstract (English)

There now exists many popular object detectors based on deep learning that can analyze images and extract locations and class labels for occurrences of objects. For image time series (i.e., video or sequences of stills), tracking objects over time and preserving object identity can help to improve object detection performance, and is necessary for many downstream tasks, including classifying and predicting behaviors, and estimating total abundances. Here we present yasmot, a lightweight and flexible object tracker that can process the output from popular object detectors and track objects over time from either monoscopic or stereoscopic camera configurations. In addition, it includes functionality to generate consensus detections from ensembles of object detectors.

多目标追踪双目视觉轻量模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。