arXiv:2511.06716cs.CVcs.AI2025-11

用Mamba模型提升视频镜面检测的准确性和鲁棒性

MirrorMamba: Towards Scalable and Robust Mirror Detection in Videos

  • 融合深度、对应关系和光流多线索,提升复杂场景适应力
  • 基于Mamba的提取器实现线性计算复杂度,捕捉全局对应关系
  • 首次将Mamba架构应用于镜面检测,适合高精度视觉任务

视频镜面检测受到广泛关注,但现有方法性能和鲁棒性有限。这些方法通常过度依赖单一不可靠的动态特征,且基于感受野受限的CNN或具有二次计算复杂度的Transformer。为此,我们提出一种高效且可扩展的视频镜面检测方法——MirrorMamba。该方法融合感知深度、对应关系和光流等多种线索,以适应多样化条件。我们引入基于Mamba的多方向对应关系提取器,利用Mamba空间状态模型的全局感受野与线性计算复杂度,有效捕捉对应关系特性。此外,设计了基于Mamba的逐层边界强化解码器,解决模糊深度图导致的边界不清晰问题。本工作首次成功将Mamba架构应用于镜面检测领域。大量实验表明,该方法在基准数据集上优于现有最先进方法;尤其在最具挑战性的基于图像的镜面检测数据集上也达到领先水平,验证了其鲁棒性与泛化能力。

原文摘要 · Abstract (English)

Video mirror detection has received significant research attention, yet existing methods suffer from limited performance and robustness. These approaches often over-rely on single, unreliable dynamic features, and are typically built on CNNs with limited receptive fields or Transformers with quadratic computational complexity. To address these limitations, we propose a new effective and scalable video mirror detection method, called MirrorMamba. Our approach leverages multiple cues to adapt to diverse conditions, incorporating perceived depth, correspondence and optical. We also introduce an innovative Mamba-based Multidirection Correspondence Extractor, which benefits from the global receptive field and linear complexity of the emerging Mamba spatial state model to effectively capture correspondence properties. Additionally, we design a Mamba-based layer-wise boundary enforcement decoder to resolve the unclear boundary caused by the blurred depth map. Notably, this work marks the first successful application of the Mamba-based architecture in the field of mirror detection. Extensive experiments demonstrate that our method outperforms existing state-of-the-art approaches for video mirror detection on the benchmark datasets. Furthermore, on the most challenging and representative image-based mirror detection dataset, our approach achieves state-of-the-art performance, proving its robustness and generalizability.

镜面检测Mamba视频理解多模态融合

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。