arXiv:2507.05948cs.CV2025-07ICCV被引 3

用深度信息提升视频实例分割的鲁棒性,解决遮挡模糊问题。

Beyond Appearance: Geometric Cues for Robust Video Instance Segmentation

  • 引入单目深度图作为几何线索,融合到分割网络中
  • 在OVIS基准上达到56.2 AP,刷新当前最佳性能
  • 适合关注视频理解中运动与遮挡场景的研究者

视频实例分割(VIS)面临物体遮挡、运动模糊和外观变化等普遍挑战。本文通过引入单目深度估计的几何感知能力,系统研究三种融合方式:扩展深度通道(EDC)将深度图作为输入通道拼接;共享ViT(SV)设计统一的ViT主干,供深度估计与分割分支共享;深度监督(DS)利用深度预测作为辅助训练信号指导特征学习。尽管DS效果有限,但基准测试表明EDC和SV显著提升了VIS鲁棒性。使用Swin-L主干时,EDC方法在OVIS基准上达到56.2 AP,创下新纪录。本工作明确证实深度线索是实现鲁棒视频理解的关键因素。

原文摘要 · Abstract (English)

Video Instance Segmentation (VIS) fundamentally struggles with pervasive challenges including object occlusions, motion blur, and appearance variations during temporal association. To overcome these limitations, this work introduces geometric awareness to enhance VIS robustness by strategically leveraging monocular depth estimation. We systematically investigate three distinct integration paradigms. Expanding Depth Channel (EDC) method concatenates the depth map as input channel to segmentation networks; Sharing ViT (SV) designs a uniform ViT backbone, shared between depth estimation and segmentation branches; Depth Supervision (DS) makes use of depth prediction as an auxiliary training guide for feature learning. Though DS exhibits limited effectiveness, benchmark evaluations demonstrate that EDC and SV significantly enhance the robustness of VIS. When with Swin-L backbone, our EDC method gets 56.2 AP, which sets a new state-of-the-art result on OVIS benchmark. This work conclusively establishes depth cues as critical enablers for robust video understanding.

视频分割几何线索深度估计鲁棒性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。