arXiv:2504.09507cs.CV2025-04

针对复杂场景优化视频对象分割,3次获挑战赛第三名。

FVOS for MOSE Track of 4th PVUW Challenge: 3rd Place Solution

  • 微调现有方法并结合形态学后处理提升精度。
  • 验证集与测试集得分分别为76.81%和83.92%。
  • 适合追求高精度视频分割的工程应用者。

视频对象分割(VOS)是计算机视觉中基础且极具挑战性的任务,应用广泛。现有方法多依赖时空记忆网络提取帧级特征,在常用数据集上表现良好,但在复杂真实场景中仍显不足。本文针对此问题,提出细调VOS(FVOS),通过定制化训练优化现有方法在特定数据集上的表现。同时引入形态学后处理策略,缓解单模型预测中相邻物体间间隙过大的问题。最后,采用基于投票的多尺度融合方法生成最终结果。该方案在2025年第四届PVUW挑战赛MOSE赛道的验证与测试阶段分别取得J&F分数76.81%和83.92%,荣获总排名第三。

原文摘要 · Abstract (English)

Video Object Segmentation (VOS) is one of the most fundamental and challenging tasks in computer vision and has a wide range of applications. Most existing methods rely on spatiotemporal memory networks to extract frame-level features and have achieved promising results on commonly used datasets. However, these methods often struggle in more complex real-world scenarios. This paper addresses this issue, aiming to achieve accurate segmentation of video objects in challenging scenes. We propose fine-tuning VOS (FVOS), optimizing existing methods for specific datasets through tailored training. Additionally, we introduce a morphological post-processing strategy to address the issue of excessively large gaps between adjacent objects in single-model predictions. Finally, we apply a voting-based fusion method on multi-scale segmentation results to generate the final output. Our approach achieves J&F scores of 76.81% and 83.92% during the validation and testing stages, respectively, securing third place overall in the MOSE Track of the 4th PVUW challenge 2025.

视频分割目标检测模型优化挑战赛

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。