用语义引导聚类提升图像令牌的语义表达能力。
SGC-VQGAN: Towards Complex Scene Representation via Semantic Guided Clustering Codebook
- 通过语义一致性学习构建时空一致的语义码本
- 在重建质量和下游任务上达到当前最优表现
- 无需额外参数训练,可直接用于下游任务
向量量化(VQ)通过离散码本表示实现特征的确定性学习。近期工作利用视觉分词器对视觉区域进行离散化,以实现自监督表征学习。然而,这些分词器存在语义缺失的问题,因其仅基于像素重构的预训练任务生成,缺乏语义信息。此外,码本分布不均和码本坍缩等问题会因码本利用率低而影响性能。为此,我们提出SGC-VQGAN,采用语义在线聚类方法,通过一致语义学习增强令牌语义。利用分割模型的推理结果,构建时空一致的语义码本,缓解码本坍缩与语义不平衡问题。提出的金字塔特征学习管道整合多层级特征,同时捕捉图像细节与语义。实验表明,SGC-VQGAN在重建质量及多种下游任务中均达到最先进水平。其结构简单,无需额外参数学习,可直接应用于下游任务,具有广泛应用潜力。
原文摘要 · Abstract (English)
Vector quantization (VQ) is a method for deterministically learning features through discrete codebook representations. Recent works have utilized visual tokenizers to discretize visual regions for self-supervised representation learning. However, a notable limitation of these tokenizers is lack of semantics, as they are derived solely from the pretext task of reconstructing raw image pixels in an auto-encoder paradigm. Additionally, issues like imbalanced codebook distribution and codebook collapse can adversely impact performance due to inefficient codebook utilization. To address these challenges, We introduce SGC-VQGAN through Semantic Online Clustering method to enhance token semantics through Consistent Semantic Learning. Utilizing inference results from segmentation model , our approach constructs a temporospatially consistent semantic codebook, addressing issues of codebook collapse and imbalanced token semantics. Our proposed Pyramid Feature Learning pipeline integrates multi-level features to capture both image details and semantics simultaneously. As a result, SGC-VQGAN achieves SOTA performance in both reconstruction quality and various downstream tasks. Its simplicity, requiring no additional parameter learning, enables its direct application in downstream tasks, presenting significant potential.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。