融合深度与彩色信息,用多存储特征记忆提升长时视频目标分割精度。
RGB-D Video Object Segmentation via Enhanced Multi-store Feature Memory
- 分层模态选择与融合,自适应结合彩色与深度特征。
- 引入SAM模型精修分割结果,有效缓解长时预测中的目标漂移。
- 适合需要高精度视频分割的机器人、自动驾驶场景使用。
RGB-D视频目标分割旨在结合彩色图像的细粒度纹理信息与深度图的空间几何线索,以提升分割性能。然而,现有方法未能充分挖掘跨模态信息,且在长期预测中易出现目标漂移。本文提出一种基于多存储特征记忆的RGB-D VOS方法,设计了分层模态选择与融合机制,自适应整合双模态特征;同时构建分割精修模块,利用分割任意模型(SAM)对分割掩码进行优化,增强记忆一致性,指导后续任务。通过时空嵌入与模态嵌入,将混合提示与融合图像输入SAM,充分发挥其在RGB-D VOS中的潜力。实验表明,该方法在最新RGB-D VOS基准上达到领先性能。
原文摘要 · Abstract (English)
The RGB-Depth (RGB-D) Video Object Segmentation (VOS) aims to integrate the fine-grained texture information of RGB with the spatial geometric clues of depth modality, boosting the performance of segmentation. However, off-the-shelf RGB-D segmentation methods fail to fully explore cross-modal information and suffer from object drift during long-term prediction. In this paper, we propose a novel RGB-D VOS method via multi-store feature memory for robust segmentation. Specifically, we design the hierarchical modality selection and fusion, which adaptively combines features from both modalities. Additionally, we develop a segmentation refinement module that effectively utilizes the Segmentation Anything Model (SAM) to refine the segmentation mask, ensuring more reliable results as memory to guide subsequent segmentation tasks. By leveraging spatio-temporal embedding and modality embedding, mixed prompts and fused images are fed into SAM to unleash its potential in RGB-D VOS. Experimental results show that the proposed method achieves state-of-the-art performance on the latest RGB-D VOS benchmark.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。