arXiv:2509.06422cs.CV2025-09

提出新方法提升视频伪装目标检测精度与泛化能力

Phantom-Insight: Adaptive Multi-cue Fusion for Video Camouflaged Object Detection with Multimodal LLM

  • 融合时空线索与多模态大模型动态生成多组提示
  • 在MoCA-Mask上达最新性能,对未见伪装物体有强泛化性
  • 适合研究视频目标分割与智能感知的开发者

视频伪装目标检测(VCOD)因环境动态变化而困难。现有方法存在两大问题:基于SAM的方法因模型冻结难以分离伪装目标边缘;基于MLLM的方法因大语言模型将前景与背景混合导致目标区分度差。为此,我们提出基于SAM与多模态大模型的新型VCOD方法Phantom-Insight。通过引入时序与空间线索,利用大模型进行特征融合以增强信息密度;设计动态前景视觉标记评分模块与提示网络,自适应引导并微调SAM模型,使其能捕捉细微纹理。进一步提出解耦式前景-背景学习策略,分别生成前景与背景提示并独立训练,使视觉标记能分别整合两类信息,从而更精准分割视频中的伪装目标。在MoCA-Mask数据集上的实验表明,Phantom-Insight在各项指标上均达到领先水平。此外,在CAD2016数据集上对未见过的伪装目标仍具良好检测能力,验证了其强大泛化性。

原文摘要 · Abstract (English)

Video camouflaged object detection (VCOD) is challenging due to dynamic environments. Existing methods face two main issues: (1) SAM-based methods struggle to separate camouflaged object edges due to model freezing, and (2) MLLM-based methods suffer from poor object separability as large language models merge foreground and background. To address these issues, we propose a novel VCOD method based on SAM and MLLM, called Phantom-Insight. To enhance the separability of object edge details, we represent video sequences with temporal and spatial clues and perform feature fusion via LLM to increase information density. Next, multiple cues are generated through the dynamic foreground visual token scoring module and the prompt network to adaptively guide and fine-tune the SAM model, enabling it to adapt to subtle textures. To enhance the separability of objects and background, we propose a decoupled foreground-background learning strategy. By generating foreground and background cues separately and performing decoupled training, the visual token can effectively integrate foreground and background information independently, enabling SAM to more accurately segment camouflaged objects in the video. Experiments on the MoCA-Mask dataset show that Phantom-Insight achieves state-of-the-art performance across various metrics. Additionally, its ability to detect unseen camouflaged objects on the CAD2016 dataset highlights its strong generalization ability.

视频检测伪装目标多模态SAM

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。