arXiv:2502.13875cs.CVcs.AI2025-02被引 3

提出MEX模块,让跟踪模型在4GB显存下高效运行

MEX: Memory-efficient Approach to Referring Multi-Object Tracking

  • 引入轻量级跨模态记忆模块,可插拔提升现有追踪器
  • 在Refer-KITTI上显著提升HOTA指标,推理速度更快
  • 适合资源受限场景,尤其适合部署在低配设备

参照式多目标跟踪(RMOT)是计算机视觉与自然语言处理交叉领域的新方向,通过文本描述识别和追踪物体,比传统方法更直观。尽管已有多种方法被提出,但多数需端到端训练整个网络。其中iKUN表现突出,我们在此基础上进一步探索并优化其流程。本文提出一种名为MEX的内存高效跨模态模块,可直接应用于现成追踪器如iKUN,实现显著架构改进。实验表明,该方法在仅4GB显存的单块GPU上推理时依然高效。在多个基准测试中,特别是包含多样化自动驾驶场景与语言描述的Refer-KITTI数据集上,该方法显著提升了HOTA指标,同时降低内存占用并加快处理速度。

原文摘要 · Abstract (English)

Referring Multi-Object Tracking (RMOT) is a relatively new concept that has rapidly gained traction as a promising research direction at the intersection of computer vision and natural language processing. Unlike traditional multi-object tracking, RMOT identifies and tracks objects and incorporates textual descriptions for object class names, making the approach more intuitive. Various techniques have been proposed to address this challenging problem; however, most require the training of the entire network due to their end-to-end nature. Among these methods, iKUN has emerged as a particularly promising solution. Therefore, we further explore its pipeline and enhance its performance. In this paper, we introduce a practical module dubbed Memory-Efficient Cross-modality -- MEX. This memory-efficient technique can be directly applied to off-the-shelf trackers like iKUN, resulting in significant architectural improvements. Our method proves effective during inference on a single GPU with 4 GB of memory. Among the various benchmarks, the Refer-KITTI dataset, which offers diverse autonomous driving scenes with relevant language expressions, is particularly useful for studying this problem. Empirically, our method demonstrates effectiveness and efficiency regarding HOTA tracking scores, substantially improving memory allocation and processing speed.

多目标跟踪跨模态内存优化轻量化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。