通过自适应伪标签优化,实现复杂视频目标分割的领先性能
STSeg-Complex Video Object Segmentation: The 1st Solution for 4th PVUW MOSE Challenge
- 融合SAM2与TMO模型,并在MOSE数据集上微调
- 推理阶段采用自适应伪标签引导优化,测试集得分87.26%
- 适合研究复杂场景下视频目标分割的开发者参考
复杂场景中的视频对象分割极具挑战性,MOSE数据集对这一领域的发展起到了重要推动作用。本文详细介绍了imaplus团队提出的STSeg解决方案。通过在MOSE数据集上对SAM2和无监督模型TMO进行微调,STSeg在处理复杂物体运动和长视频序列方面表现出显著优势。推理阶段采用自适应伪标签引导的模型优化流程,智能选择适用于每段视频的模型。通过模型微调与推理阶段的自适应优化,STSeg在2025年第四届PVUW挑战赛MOSE赛道测试集上取得了87.26%的J&F得分,获得第一名,推动了复杂场景下视频对象分割技术的发展。
原文摘要 · Abstract (English)
Segmentation of video objects in complex scenarios is highly challenging, and the MOSE dataset has significantly contributed to the development of this field. This technical report details the STSeg solution proposed by the "imaplus" team.By finetuning SAM2 and the unsupervised model TMO on the MOSE dataset, the STSeg solution demonstrates remarkable advantages in handling complex object motions and long-video sequences. In the inference phase, an Adaptive Pseudo-labels Guided Model Refinement Pipeline is adopted to intelligently select appropriate models for processing each video. Through finetuning the models and employing the Adaptive Pseudo-labels Guided Model Refinement Pipeline in the inference phase, the STSeg solution achieved a J&F score of 87.26% on the test set of the 2025 4th PVUW Challenge MOSE Track, securing the 1st place and advancing the technology for video object segmentation in complex scenarios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。