arXiv:2509.14060cs.CV2025-09

用视觉语义增强提升低画质视频的多目标追踪效果

VSE-MOT: Multi-Object Tracking in Low-Quality Video Scenes Guided by Visual Semantic Enhancement

  • 引入视觉语言模型提取全局语义信息,融合查询向量增强特征
  • 在真实低质视频中性能提升8%至20%,常规场景也保持稳定
  • 适合低画质复杂环境下的追踪任务,尤其适用于实际应用

当前多目标追踪(MOT)算法通常忽略低质量视频中的固有问题,导致在现实图像退化场景下性能显著下降。为此,本文提出一种视觉语义增强引导的多目标追踪框架(VSE-MOT)。设计三分支结构,利用视觉语言模型提取图像全局视觉语义信息,并与查询向量融合。为进一步提升语义信息利用效率,引入多目标追踪适配器(MOT-Adapter)和视觉语义融合模块(VSFM),分别实现语义信息的任务适配与特征融合优化。大量实验验证了该方法在真实低质量视频场景下的有效性与优越性,追踪性能较现有方法提升约8%至20%,同时在常规场景中仍保持稳健表现。

原文摘要 · Abstract (English)

Current multi-object tracking (MOT) algorithms typically overlook issues inherent in low-quality videos, leading to significant degradation in tracking performance when confronted with real-world image deterioration. Therefore, advancing the application of MOT algorithms in real-world low-quality video scenarios represents a critical and meaningful endeavor. To address the challenges posed by low-quality scenarios, inspired by vision-language models, this paper proposes a Visual Semantic Enhancement-guided Multi-Object Tracking framework (VSE-MOT). Specifically, we first design a tri-branch architecture that leverages a vision-language model to extract global visual semantic information from images and fuse it with query vectors. Subsequently, to further enhance the utilization of visual semantic information, we introduce the Multi-Object Tracking Adapter (MOT-Adapter) and the Visual Semantic Fusion Module (VSFM). The MOT-Adapter adapts the extracted global visual semantic information to suit multi-object tracking tasks, while the VSFM improves the efficacy of feature fusion. Through extensive experiments, we validate the effectiveness and superiority of the proposed method in real-world low-quality video scenarios. Its tracking performance metrics outperform those of existing methods by approximately 8% to 20%, while maintaining robust performance in conventional scenarios.

多目标追踪低质量视频视觉语义视觉语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。