arXiv:2508.04211cs.CV2025-08ICCV被引 3

破解开放词汇分割的性能瓶颈,揭示模型失败根源。

What Holds Back Open-Vocabulary Segmentation?

  • 利用真实标签信息分离并诊断性能瓶颈
  • 发现两年来性能停滞的关键原因
  • 为未来研究提供可落地的突破方向

标准分割设置无法实现对训练类别外概念的识别。开放词汇方法通过在数十亿图像-标题对上进行语言-图像预训练,有望弥补这一差距。然而,我们观察到该承诺并未实现,因为存在多个导致性能停滞近两年的瓶颈。本文提出新型基准组件,借助真实标签信息识别并解耦这些瓶颈。验证实验揭示了开放词汇模型失败的深层原因,提供了重要实证发现,为未来研究指明了关键突破路径。

原文摘要 · Abstract (English)

Standard segmentation setups are unable to deliver models that can recognize concepts outside the training taxonomy. Open-vocabulary approaches promise to close this gap through language-image pretraining on billions of image-caption pairs. Unfortunately, we observe that the promise is not delivered due to several bottlenecks that have caused the performance to plateau for almost two years. This paper proposes novel oracle components that identify and decouple these bottlenecks by taking advantage of the groundtruth information. The presented validation experiments deliver important empirical findings that provide a deeper insight into the failures of open-vocabulary models and suggest prominent approaches to unlock the future research.

开放词汇语义分割模型瓶颈

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。