arXiv:2508.18050cs.CV2025-08

用类人思维推理实现隐身物体精准分割,零样本效果领先

ArgusCogito: Chain-of-Thought for Cross-Modal Synergy and Omnidirectional Reasoning in Camouflaged Object Segmentation

  • 通过跨模态融合与多向推理构建认知先验,提升整体理解力
  • 在四个隐身物体分割数据集上达到最新最优性能,平均DICE超75%
  • 适合需要强泛化能力的医学图像分割等高精度场景应用

隐身物体分割(COS)因目标与背景高度相似而极具挑战,需模型具备超越表面线索的深层全局理解。现有方法受限于浅层特征表示、推理机制不足及跨模态整合弱,常出现目标分离不全、分割不精确等问题。受‘百眼巨人’整体观察能力与多向专注启发,我们提出ArgusCogito,一种基于视觉-语言模型的零样本链式思维框架,融合跨模态协同与全向推理。该框架包含三个认知驱动阶段:(1) 推测:通过RGB、深度图与语义图的跨模态融合进行全局推理,建立强认知先验,增强目标-背景区分;(2) 聚焦:基于推测阶段的语义先验,执行多向注意力扫描与聚焦推理,精确定位目标并优化关注区域;(3) 雕刻:在聚焦区域内迭代生成密集正负提示,融合跨模态信息,逐步塑造高保真分割掩码,模拟高强度审视过程。在四个挑战性COS基准和三个医学图像分割(MIS)基准上的广泛评估表明,ArgusCogito实现最先进(SOTA)性能,验证了其卓越有效性、优异泛化能力与鲁棒性。

原文摘要 · Abstract (English)

Camouflaged Object Segmentation (COS) poses a significant challenge due to the intrinsic high similarity between targets and backgrounds, demanding models capable of profound holistic understanding beyond superficial cues. Prevailing methods, often limited by shallow feature representation, inadequate reasoning mechanisms, and weak cross-modal integration, struggle to achieve this depth of cognition, resulting in prevalent issues like incomplete target separation and imprecise segmentation. Inspired by the perceptual strategy of the Hundred-eyed Giant-emphasizing holistic observation, omnidirectional focus, and intensive scrutiny-we introduce ArgusCogito, a novel zero-shot, chain-of-thought framework underpinned by cross-modal synergy and omnidirectional reasoning within Vision-Language Models (VLMs). ArgusCogito orchestrates three cognitively-inspired stages: (1) Conjecture: Constructs a strong cognitive prior through global reasoning with cross-modal fusion (RGB, depth, semantic maps), enabling holistic scene understanding and enhanced target-background disambiguation. (2) Focus: Performs omnidirectional, attention-driven scanning and focused reasoning, guided by semantic priors from Conjecture, enabling precise target localization and region-of-interest refinement. (3) Sculpting: Progressively sculpts high-fidelity segmentation masks by integrating cross-modal information and iteratively generating dense positive/negative point prompts within focused regions, emulating Argus' intensive scrutiny. Extensive evaluations on four challenging COS benchmarks and three Medical Image Segmentation (MIS) benchmarks demonstrate that ArgusCogito achieves state-of-the-art (SOTA) performance, validating the framework's exceptional efficacy, superior generalization capability, and robustness.

隐身分割链式推理跨模态融合医学图像

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。