无需标注,通过时间一致性提升视频目标检测与分割精度
VVitCutLER: Towards Unsupervised Object Detection and Segmentation in Videos

- 利用跨帧区域一致性生成稳定伪标签,抑制误差累积
- 在DAVIS、YouTube-VOS等数据集上实现领先性能,显著减少伪标签闪烁
- 适合研究无监督视频理解或需要鲁棒性分割的场景
无监督像素级视频理解在真实场景中仍具挑战,运动模糊、遮挡和快速物体动态常导致时间漂移和闪烁伪标签。我们提出VVitCutLER,一种无监督视频目标检测与实例分割框架,通过时间一致性提升伪标签质量。核心贡献是VitCut,一种具备时序稳定性的伪标签生成器,通过跨帧区域一致性减少场退化下的误差累积;同时采用蒸馏解码器实现有效的实例掩码预测。基于VitCut,VVitCutLER进一步融合跨帧特征聚合,增强视频级鲁棒性。在标准视频基准(如DAVIS、YouTube-VOS)上的大量实验表明,该方法显著提升检测与分割性能,同时降低时间不稳定性。结果凸显了时间一致性监督对鲁棒像素级视频理解的重要性。
原文摘要 · Abstract (English)
Unsupervised pixel-level video understanding remains challenging in real-world scenarios, where motion blur, occlusion, and fast object dynamics often cause temporal drift and flickering pseudo-labels.We propose VVitCutLER, an unsupervised framework for video object detection and instance segmentation, which improves the quality of pseudo-labels through temporal consistency. Our core contribution is VitCut, a temporarily stable pseudo-label generator that reduces error accumulation during field degradation through cross-frame region consistency. Meanwhile, VitCut uses a distillation decoder to achieve effective instance mask prediction. Then, based on VitCut, VVitCutLER further integrates cross-frame feature aggregation to enhance video-level robustness. Extensive experiments on standard video benchmarks demonstrate that VVitCutLER significantly improves detection and segmentation performance while reducing temporal instability. These results highlight the importance of temporally consistent supervision for robust pixel-level video understanding.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。