用信息论方法雕刻表征空间,提升未知类别发现能力。
InfoSculpt: Sculpting the Latent Space for Generalized Category Discovery
- 基于信息瓶颈原理,双目标优化分离类别信号与噪声
- 在8个基准上显著优于现有方法,实现更鲁棒的分类
- 适合开放世界场景下的无监督类别发现任务
广义类别发现(GCD)旨在对大规模无标签数据中的已知和未知类别实例进行分类,是现实世界开放环境应用的关键挑战。现有方法多依赖伪标签或两阶段聚类,缺乏明确机制来解耦类别定义信号与实例特异性噪声。本文从信息论视角重构GCD问题,基于信息瓶颈(IB)原则提出InfoSculpt框架,通过最小化双重条件互信息(CMI)目标系统性地雕刻表示空间。该框架在有标签数据上施加类别级CMI以学习紧凑判别特征,在全部数据上施加实例级CMI以压缩增强带来的噪声并提取不变特征。两者在不同尺度协同作用,生成解耦且鲁棒的潜在空间,保留类别信息同时剔除噪声细节。在8个基准上的大量实验验证了该信息论方法的有效性。
原文摘要 · Abstract (English)
Generalized Category Discovery (GCD) aims to classify instances from both known and novel categories within a large-scale unlabeled dataset, a critical yet challenging task for real-world, open-world applications. However, existing methods often rely on pseudo-labeling, or two-stage clustering, which lack a principled mechanism to explicitly disentangle essential, category-defining signals from instance-specific noise. In this paper, we address this fundamental limitation by re-framing GCD from an information-theoretic perspective, grounded in the Information Bottleneck (IB) principle. We introduce InfoSculpt, a novel framework that systematically sculpts the representation space by minimizing a dual Conditional Mutual Information (CMI) objective. InfoSculpt uniquely combines a Category-Level CMI on labeled data to learn compact and discriminative representations for known classes, and a complementary Instance-Level CMI on all data to distill invariant features by compressing augmentation-induced noise. These two objectives work synergistically at different scales to produce a disentangled and robust latent space where categorical information is preserved while noisy, instance-specific details are discarded. Extensive experiments on 8 benchmarks demonstrate that InfoSculpt validating the effectiveness of our information-theoretic approach.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。