arXiv:2501.09688cs.CV2025-01CVPR被引 7

通过分层成本聚合与结构引导,提升未知类别物体细粒度分割精度。

Fine-Grained Image-Text Correspondence with Cost Aggregation for Open-Vocabulary Part Segmentation

  • 分设物体与部件级成本聚合,精准定位细粒度部分
  • 在3个数据集上超越现有方法,显著提升对未见类别的泛化能力
  • 适合需要细粒度语义分割的开放词汇场景

开放词汇部件分割(OVPS)旨在识别未见类别中的细粒度部件。本文指出两大挑战:部件级图像-文本对齐困难,以及部件分割缺乏结构理解。为此提出PartCATSeg框架,融合对象感知的部件级成本聚合、组合损失及来自DINO的结构引导。采用解耦成本聚合策略,分别处理对象与部件级成本,提升部件分割精度;引入组合损失以更好建模部件-对象关系,弥补部件标注不足;利用DINO特征提供结构引导,改善边界划分与部件间关系理解。在Pascal-Part-116、ADE20K-Part-234和PartImageNet三个数据集上的实验表明,该方法显著优于现有先进方法,为未见部件类别的鲁棒泛化设立了新基准。

原文摘要 · Abstract (English)

Open-Vocabulary Part Segmentation (OVPS) is an emerging field for recognizing fine-grained parts in unseen categories. We identify two primary challenges in OVPS: (1) the difficulty in aligning part-level image-text correspondence, and (2) the lack of structural understanding in segmenting object parts. To address these issues, we propose PartCATSeg, a novel framework that integrates object-aware part-level cost aggregation, compositional loss, and structural guidance from DINO. Our approach employs a disentangled cost aggregation strategy that handles object and part-level costs separately, enhancing the precision of part-level segmentation. We also introduce a compositional loss to better capture part-object relationships, compensating for the limited part annotations. Additionally, structural guidance from DINO features improves boundary delineation and inter-part understanding. Extensive experiments on Pascal-Part-116, ADE20K-Part-234, and PartImageNet datasets demonstrate that our method significantly outperforms state-of-the-art approaches, setting a new baseline for robust generalization to unseen part categories.

细粒度分割开放词汇图文对齐结构引导

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。