arXiv:2410.01341cs.CV2024-10中稿 · IEEE Transactions …被引 7

通过解耦认知迁移提升文本监督下的第一视角语义分割性能

Cognition Transferring and Decoupling for Text-supervised Egocentric Semantic Segmentation

  • 利用图文关联学习第一视角中穿戴者与物体的关系
  • 在四个基准上显著超越现有方法,实现更精准的前景背景区分
  • 适合研究第一视角视觉理解与跨模态学习的开发者

本文提出一种新型文本监督的第一视角语义分割(TESS)任务,旨在仅用图像级文本标签对第一视角图像进行像素级分类。该任务面临密集的穿戴者-物体关系及物体间干扰问题。现有基于冻结CLIP模型的方法因缺乏对关系的敏感性而表现不佳。为此,我们提出认知迁移与解耦网络(CTDN),首先通过图文关联学习第一视角中的穿戴者-物体关系;设计认知迁移模块(CTM)从大规模预训练模型中提炼认知知识,以识别具有多样语义的第一视角物体;基于迁移的认知,前景-背景解耦模块(FDM)显式分离视觉表征,缓解因前景-背景干扰物体导致的误激活区域。在四个TESS基准上的实验表明,所提方法显著优于多种近期方法。代码将公开于https://github.com/ZhaofengSHI/CTDN。

原文摘要 · Abstract (English)

In this paper, we explore a novel Text-supervised Egocentic Semantic Segmentation (TESS) task that aims to assign pixel-level categories to egocentric images weakly supervised by texts from image-level labels. In this task with prospective potential, the egocentric scenes contain dense wearer-object relations and inter-object interference. However, most recent third-view methods leverage the frozen Contrastive Language-Image Pre-training (CLIP) model, which is pre-trained on the semantic-oriented third-view data and lapses in the egocentric view due to the ``relation insensitive" problem. Hence, we propose a Cognition Transferring and Decoupling Network (CTDN) that first learns the egocentric wearer-object relations via correlating the image and text. Besides, a Cognition Transferring Module (CTM) is developed to distill the cognitive knowledge from the large-scale pre-trained model to our model for recognizing egocentric objects with various semantics. Based on the transferred cognition, the Foreground-background Decoupling Module (FDM) disentangles the visual representations to explicitly discriminate the foreground and background regions to mitigate false activation areas caused by foreground-background interferential objects during egocentric relation learning. Extensive experiments on four TESS benchmarks demonstrate the effectiveness of our approach, which outperforms many recent related methods by a large margin. Code will be available at https://github.com/ZhaofengSHI/CTDN.

第一视角语义分割跨模态认知迁移

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。