MASSeg提升复杂视频目标分割性能,应对小目标与遮挡挑战。
MASSeg : 2nd Technical Report for 4th PVUW MOSE Track
- 融合帧间一致与不一致数据增强,提升模型鲁棒性。
- 推理时采用掩码缩放策略,适应不同大小和遮挡程度的目标。
- 在MOSE+数据集上取得J=0.8250、F=0.9007的优异表现。
复杂视频目标分割在小目标识别、遮挡处理和动态场景建模方面仍面临重大挑战。本报告介绍我们的解决方案,该方案在CVPR 2025 PVUW挑战赛的MOSE赛道中位列第二。基于现有分割框架,我们提出改进模型MASSeg,并构建包含典型遮挡、杂乱背景和小目标实例的增强数据集MOSE+。训练阶段,采用帧间一致与不一致数据增强组合策略,提升模型泛化能力;推理阶段,设计掩码输出缩放策略,更好适应不同目标尺寸与遮挡水平。最终,MASSeg在MOSE测试集上实现J得分0.8250、F得分0.9007,以及J&F得分0.8628。
原文摘要 · Abstract (English)
Complex video object segmentation continues to face significant challenges in small object recognition, occlusion handling, and dynamic scene modeling. This report presents our solution, which ranked second in the MOSE track of CVPR 2025 PVUW Challenge. Based on an existing segmentation framework, we propose an improved model named MASSeg for complex video object segmentation, and construct an enhanced dataset, MOSE+, which includes typical scenarios with occlusions, cluttered backgrounds, and small target instances. During training, we incorporate a combination of inter-frame consistent and inconsistent data augmentation strategies to improve robustness and generalization. During inference, we design a mask output scaling strategy to better adapt to varying object sizes and occlusion levels. As a result, MASSeg achieves a J score of 0.8250, F score of 0.9007, and a J&F score of 0.8628 on the MOSE test set.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。