arXiv:2604.00395cs.CV2026-04

用追踪增强提示提升小目标和语义主导目标的视频分割效果

Advancing Complex Video Object Segmentation via Tracking-Enhanced Prompt: The 1st Winner for 5th PVUW MOSE Challenge

  • 引入外部追踪模型与多模态大模型生成追踪增强提示
  • 在PVUW挑战赛中取得56.91%的测试集分数排名第一
  • 无需训练,可直接提升SAM3对复杂目标的理解能力

在复杂视频目标分割任务中,研究者需在杂乱环境中跟踪并分割特定目标,这对模型的目标理解与环境适应能力提出了严格考验。尽管当前最先进方法SAM3在常规目标上表现出色,但在微小目标和语义主导目标上表现不佳,根源在于其对这类目标理解不足。为此,我们提出TEP:通过追踪增强提示推进复杂视频目标分割。作为一项无需训练的方法,TEP利用外部追踪模型与多模态大语言模型生成追踪增强提示,缓解SAM3在理解此类挑战性目标时的困难。该方法在PVUW Challenge 2026:复杂视频目标分割赛道的测试集上取得56.91%的成绩,位列第一。

原文摘要 · Abstract (English)

In the Complex Video Object Segmentation task, researchers are required to track and segment specific targets within cluttered environments, which rigorously tests a method's capability for target comprehension and environmental adaptability. Although SAM3, the current state-of-the-art solution, exhibits unparalleled segmentation performance and robustness on conventional targets, it underperforms on tiny and semantic-dominated objects. The root cause of this limitation lies in SAM3's insufficient comprehension of these specific target types. To address this issue, we propose TEP: Advancing Complex Video Object Segmentation via Tracking-Enhanced Prompts. As a training-free approach, TEP leverages external tracking models and Multimodal Large Language Models to introduce tracking-enhanced prompts, thereby alleviating the difficulty SAM3 faces in understanding these challenging targets. Our method achieved first place (56.91%) on the test set of the PVUW Challenge 2026: Complex Video Object Segmentation Track.

视频分割追踪增强SAM3零样本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。