arXiv:2601.15931cs.AIcs.LG2026-01被引 1

提出因果干预与符号先验结合的新框架,提升文本行人搜索在复杂场景下的鲁棒性。

ICON: Invariant Counterfactual Optimization with Neuro-Symbolic Priors for Text-Based Person Search

  • 引入规则引导的空间干预,消除边界框噪声带来的几何偏差。
  • 通过反事实背景替换,使模型忽略环境干扰,实现环境独立性。
  • 融合神经符号拓扑对齐,确保特征激活符合人体结构逻辑,适合复杂监控场景使用。

文本驱动的行人搜索(TBPS)在真实监控场景中具有独特价值,连接视觉感知与语言理解,但当前基于预训练模型的方法在开放世界复杂场景下泛化能力差。其依赖‘被动观察’导致多重虚假相关和空间语义错位,难以应对分布偏移。为此,本文提出ICON(不变反事实优化与神经符号先验),融合因果与拓扑先验:首先,规则引导的空间干预严格惩罚对边界框噪声的敏感性,强制切断位置捷径,实现几何不变性;其次,通过语义驱动的背景移植实现反事实上下文解耦,迫使模型忽略背景干扰,获得环境独立性;再次,采用显著性驱动的语义正则化与自适应掩码,缓解局部显著性偏差,保证语义完整性;最后,利用神经符号拓扑对齐,以神经符号先验约束特征匹配,确保激活区域在拓扑上符合人体结构逻辑。实验表明,ICON不仅在标准基准上保持领先性能,还在遮挡、背景干扰和定位噪声下表现出卓越鲁棒性。该方法推动领域从拟合统计共现转向学习因果不变性。

原文摘要 · Abstract (English)

Text-Based Person Search (TBPS) holds unique value in real-world surveillance bridging visual perception and language understanding, yet current paradigms utilizing pre-training models often fail to transfer effectively to complex open-world scenarios. The reliance on "Passive Observation" leads to multifaceted spurious correlations and spatial semantic misalignment, causing a lack of robustness against distribution shifts. To fundamentally resolve these defects, this paper proposes ICON (Invariant Counterfactual Optimization with Neuro-symbolic priors), a framework integrating causal and topological priors. First, we introduce Rule-Guided Spatial Intervention to strictly penalize sensitivity to bounding box noise, forcibly severing location shortcuts to achieve geometric invariance. Second, Counterfactual Context Disentanglement is implemented via semantic-driven background transplantation, compelling the model to ignore background interference for environmental independence. Then, we employ Saliency-Driven Semantic Regularization with adaptive masking to resolve local saliency bias and guarantee holistic completeness. Finally, Neuro-Symbolic Topological Alignment utilizes neuro-symbolic priors to constrain feature matching, ensuring activated regions are topologically consistent with human structural logic. Experimental results demonstrate that ICON not only maintains leading performance on standard benchmarks but also exhibits exceptional robustness against occlusion, background interference, and localization noise. This approach effectively advances the field by shifting from fitting statistical co-occurrences to learning causal invariance.

文本搜索因果推理行人检索鲁棒性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。