用统一提示提升多模态伪装目标检测,适配任意新模态
Modality-Agnostic Prompt Learning for Multi-Modal Camouflaged Object Detection

- 构建跨模态通用提示,无需修改模型架构
- 在三组多模态数据上显著提升分割精度
- 适合需要快速接入新传感器的视觉系统
伪装目标检测(COD)旨在分割与复杂背景融为一体的物体,近年来通过引入额外视觉模态以利用互补信息提升鲁棒性。然而,现有方法普遍依赖特定模态的架构或定制融合策略,限制了可扩展性和跨模态泛化能力。为此,本文提出一种新框架,为分割一切模型(SAM)生成跨模态通用提示,实现对任意辅助模态的参数高效适配,并显著提升COD任务性能。具体而言,通过数据驱动的内容域与知识驱动的提示域之间的交互建模多模态学习,将任务相关线索提炼为统一提示供SAM解码。此外,引入轻量级掩码精修模块,利用细粒度提示线索校正粗略预测,获得更精确的伪装目标边界。在RGB-Depth、RGB-Thermal和RGB-Polarization基准上的大量实验验证了所提框架的有效性与泛化能力。
原文摘要 · Abstract (English)
Camouflaged Object Detection (COD) aims to segment objects that blend seamlessly into complex backgrounds, with growing interest in exploiting additional visual modalities to enhance robustness through complementary information. However, most existing approaches generally rely on modality-specific architectures or customized fusion strategies, which limit scalability and cross-modal generalization. To address this, we propose a novel framework that generates modality-agnostic multi-modal prompts for the Segment Anything Model (SAM), enabling parameter-efficient adaptation to arbitrary auxiliary modalities and significantly improving overall performance on COD tasks. Specifically, we model multi-modal learning through interactions between a data-driven content domain and a knowledge-driven prompt domain, distilling task-relevant cues into unified prompts for SAM decoding. We further introduce a lightweight Mask Refine Module to calibrate coarse predictions by incorporating fine-grained prompt cues, leading to more accurate camouflaged object boundaries. Extensive experiments on RGB-Depth, RGB-Thermal, and RGB-Polarization benchmarks validate the effectiveness and generalization of our modality-agnostic framework.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。