arXiv:2409.14343cs.CVeess.IV2024-09中稿 · ICPR2024

通过联合优化匹配与解码,提升视频物体分割精度。

Memory Matching is not Enough: Jointly Improving Memory Matching and Decoding for Video Object Segmentation

  • 引入成本感知机制和跨尺度匹配,增强记忆匹配能力
  • 在DAVIS2016/2017验证集上达到92.4%/88.1%准确率
  • 适合需要高精度视频分割的科研与工业应用

基于记忆的视频物体分割方法通过建立记忆库,在长时空跨度上建模多个物体,取得了显著性能。然而,这类方法易产生误匹配,且容易丢失关键信息,导致不同物体间混淆。本文提出一种联合优化匹配与解码阶段的方法,以缓解误匹配问题。在记忆匹配阶段,设计了成本感知机制,抑制短期记忆中的微小误差;并引入分流式跨尺度匹配,为不同物体尺度构建更广的匹配空间。在读出解码阶段,实施补偿机制,恢复匹配阶段缺失的关键信息。该方法在多个主流基准上表现优异:DAVIS 2016&2017验证集分别达92.4%和88.1%,DAVIS 2017测试集达83.9%,YouTubeVOS 2018&2019验证集分别为84.8%和84.6%。

原文摘要 · Abstract (English)

Memory-based video object segmentation methods model multiple objects over long temporal-spatial spans by establishing memory bank, which achieve the remarkable performance. However, they struggle to overcome the false matching and are prone to lose critical information, resulting in confusion among different objects. In this paper, we propose an effective approach which jointly improving the matching and decoding stages to alleviate the false matching issue.For the memory matching stage, we present a cost aware mechanism that suppresses the slight errors for short-term memory and a shunted cross-scale matching for long-term memory which establish a wide filed matching spaces for various object scales. For the readout decoding stage, we implement a compensatory mechanism aims at recovering the essential information where missing at the matching stage. Our approach achieves the outstanding performance in several popular benchmarks (i.e., DAVIS 2016&2017 Val (92.4%&88.1%), and DAVIS 2017 Test (83.9%)), and achieves 84.8%&84.6% on YouTubeVOS 2018&2019 Val.

视频分割记忆机制解码优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。