用深度信息提升视频镜面分割,自动识别模糊反射和弱边界镜面。
MirrorSAM2: Segment Mirror in Videos with Depth Perception
- 融合RGB与深度图,通过四模块协同解决镜面歧义问题。
- 在VMD和DVMD数据集上达当前最佳,小镜面与强反射下仍稳定表现。
- 首次实现SAM2零提示视频镜面分割,适合视觉系统开发与自动驾驶应用。
本文提出MirrorSAM2,首个将Segment Anything Model 2(SAM2)适配至RGB-D视频镜面分割的任务框架。针对镜面检测中的反射歧义与纹理混淆等关键挑战,引入四个定制化模块:深度畸变模块用于对齐RGB与深度图,深度引导的多尺度点提示生成器实现自动提示生成,频域细节注意力融合模块增强结构边界,以及带可学习镜面标记的镜面掩码解码器实现精细化分割。通过充分挖掘RGB与深度图的互补性,MirrorSAM2实现了完全无需提示的视频镜面分割能力。据我们所知,这是首个使SAM2具备自动视频镜面分割能力的工作。在VMD与DVMD基准测试中,该方法取得当前最优性能,即使在小镜面、弱边界及强反射等复杂条件下也表现稳健。
原文摘要 · Abstract (English)
This paper presents MirrorSAM2, the first framework that adapts Segment Anything Model 2 (SAM2) to the task of RGB-D video mirror segmentation. MirrorSAM2 addresses key challenges in mirror detection, such as reflection ambiguity and texture confusion, by introducing four tailored modules: a Depth Warping Module for RGB and depth alignment, a Depth-guided Multi-Scale Point Prompt Generator for automatic prompt generation, a Frequency Detail Attention Fusion Module to enhance structural boundaries, and a Mirror Mask Decoder with a learnable mirror token for refined segmentation. By fully leveraging the complementarity between RGB and depth, MirrorSAM2 extends SAM2's capabilities to the prompt-free setting. To our knowledge, this is the first work to enable SAM2 for automatic video mirror segmentation. Experiments on the VMD and DVMD benchmark demonstrate that MirrorSAM2 achieves SOTA performance, even under challenging conditions such as small mirrors, weak boundaries, and strong reflections.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。