arXiv:2505.06524cs.CV2025-05

通过因果提示校准提升SAM在开放词汇多实体分割中的泛化能力

Causal Prompt Calibration Guided Segment Anything Model for Open-Vocabulary Multi-Entity Segmentation

  • 基于因果分析,识别提示中的干扰因子是导致泛化差的主因
  • 提出CPC-SAM方法,通过多分布一致性约束生成仅含相关因素的因果提示
  • 轻量级提示学习器实现端到端优化,适合开放词汇分割任务

尽管分割一切模型(SAM)表现强大,但在开放词汇多实体分割(OVMS)中仍存在泛化问题。通过实证与因果分析发现:(i) 提示偏差是泛化问题的主要原因;(ii) 此偏差与提示中无关任务的生成因素密切相关,这些因素构成混杂变量并影响泛化效果。为此,我们提出一种可校准提示以消除混杂因子的方法。基于因果分析,定义最优提示应仅包含任务相关的因果因素,称为因果提示。理论分析表明,在因果多分布一致性理论基础上,可通过强制分割一致性和最优性来获得该提示。受此启发,提出CPC-SAM,一种用于SAM的因果提示校准方法,以实现准确的OVMS。其集成轻量级因果提示学习器(CaPL),先用随机标注生成多个提示以模拟不同分布,再通过CaPL在任务和实体层面强制因果多分布一致性进行重加权。为确保获取因果提示,CaPL通过最小化重加权提示上的累积分割损失进行优化,以达成一致性和最优性。采用双层优化策略交替优化CaPL与SAM,保障准确的OVMS。大量实验验证其优越性。

原文摘要 · Abstract (English)

Despite the strength of the Segment Anything Model (SAM), it struggles with generalization issues in open-vocabulary multi-entity segmentation (OVMS). Through empirical and causal analyses, we find that (i) the prompt bias is the primary cause of the generalization issues; (ii) this bias is closely tied to the task-irrelevant generating factors within the prompts, which act as confounders and affect generalization. To address the generalization issues, we aim to propose a method that can calibrate prompts to eliminate confounders for accurate OVMS. Building upon the causal analysis, we propose that the optimal prompt for OVMS should contain only task-relevant causal factors. We define it as the causal prompt, serving as the goal of calibration. Next, our theoretical analysis, grounded by causal multi-distribution consistency theory, proves that this prompt can be obtained by enforcing segmentation consistency and optimality. Inspired by this, we propose CPC-SAM, a Causal Prompt Calibration method for SAM to achieve accurate OVMS. It integrates a lightweight causal prompt learner (CaPL) into SAM to obtain causal prompts. Specifically, we first generate multiple prompts using random annotations to simulate diverse distributions and then reweight them via CaPL by enforcing causal multi-distribution consistency in both task and entity levels. To ensure obtaining causal prompts, CaPL is optimized by minimizing the cumulative segmentation loss across the reweighted prompts to achieve consistency and optimality. A bi-level optimization strategy alternates between optimizing CaPL and SAM, ensuring accurate OVMS. Extensive experiments validate its superiority.

分割一切因果推理开放词汇

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。