用因果干预提升第一视角视频目标分割的鲁棒性
Robust Egocentric Referring Video Object Segmentation via Dual-Modal Causal Intervention
- 通过双模因果干预,分别处理语言偏见和视觉混淆
- 在Ego-RVOS基准上达到当前最优性能
- 适合关注第一视角视频理解与模型可靠性研究者
第一人称指代视频目标分割(Ego-RVOS)旨在根据语言描述,在第一人称视频中分割出与人类动作直接相关的特定对象。该任务对理解第一人称人类行为至关重要。然而,由于第一人称视频固有的歧义性以及训练数据中的偏差,实现鲁棒分割仍具挑战。现有方法常因数据集中扭曲的对象-动作配对关系而学习到虚假关联,并受第一人称视角下的视觉混杂因素(如快速运动、频繁遮挡)影响。为此,我们提出因果第一人称指代分割(CERES),一种可插入的因果框架,将强大的预训练RVOS主干适配至第一人称领域。CERES实施双模因果干预:运用后门调整原则对抗由数据统计学习的语言表示偏差;利用前门调整概念,通过因果指导,智能融合语义视觉特征与几何深度信息,构建更抗第一人称失真的表示。大量实验表明,CERES在Ego-RVOS基准上达到当前最优性能,凸显了因果推理在构建更可靠的第一人称视频理解模型方面的潜力。
原文摘要 · Abstract (English)
Egocentric Referring Video Object Segmentation (Ego-RVOS) aims to segment the specific object actively involved in a human action, as described by a language query, within first-person videos. This task is critical for understanding egocentric human behavior. However, achieving such segmentation robustly is challenging due to ambiguities inherent in egocentric videos and biases present in training data. Consequently, existing methods often struggle, learning spurious correlations from skewed object-action pairings in datasets and fundamental visual confounding factors of the egocentric perspective, such as rapid motion and frequent occlusions. To address these limitations, we introduce Causal Ego-REferring Segmentation (CERES), a plug-in causal framework that adapts strong, pre-trained RVOS backbones to the egocentric domain. CERES implements dual-modal causal intervention: applying backdoor adjustment principles to counteract language representation biases learned from dataset statistics, and leveraging front-door adjustment concepts to address visual confounding by intelligently integrating semantic visual features with geometric depth information guided by causal principles, creating representations more robust to egocentric distortions. Extensive experiments demonstrate that CERES achieves state-of-the-art performance on Ego-RVOS benchmarks, highlighting the potential of applying causal reasoning to build more reliable models for broader egocentric video understanding.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。