将SAM与Mamba结合,提升光场图像显著目标检测精度
LFSamba: Marry SAM with Mamba for Light Field Salient Object Detection
- 用SAM提取多焦点特征,实现模态感知的高效特征提取
- 通过Mamba建模跨焦距切片长程依赖,捕捉隐含深度信息
- 支持弱监督学习,首次构建光场显著目标检测的涂鸦标注基线
光场相机可通过捕获的多焦点图像重建三维场景,蕴含丰富的空间几何信息,广泛应用于立体摄影、虚拟现实和机器人视觉。本文提出一种针对多焦点光场图像的前沿显著目标检测模型LFSamba,基于四大核心洞察:(a) 高效特征提取:利用SAM提取模态感知的判别性特征;(b) 切片间关系建模:借助Mamba捕捉跨多个焦距切片的长程依赖,从而提取隐含深度线索;(c) 模态间关系建模:使用Mamba融合全聚焦与多焦点图像,实现相互增强;(d) 弱监督学习能力:从现有像素级掩码数据集构建涂鸦标注数据集,建立首个光场显著目标检测的涂鸦监督基线。
原文摘要 · Abstract (English)
A light field camera can reconstruct 3D scenes using captured multi-focus images that contain rich spatial geometric information, enhancing applications in stereoscopic photography, virtual reality, and robotic vision. In this work, a state-of-the-art salient object detection model for multi-focus light field images, called LFSamba, is introduced to emphasize four main insights: (a) Efficient feature extraction, where SAM is used to extract modality-aware discriminative features; (b) Inter-slice relation modeling, leveraging Mamba to capture long-range dependencies across multiple focal slices, thus extracting implicit depth cues; (c) Inter-modal relation modeling, utilizing Mamba to integrate all-focus and multi-focus images, enabling mutual enhancement; (d) Weakly supervised learning capability, developing a scribble annotation dataset from an existing pixel-level mask dataset, establishing the first scribble-supervised baseline for light field salient object detection.https://github.com/liuzywen/LFScribble
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。