arXiv:2607.24943cs.LGastro-ph.GA2026-07被引 1

仅凭混合样本的身份,就能从无标签数据中发现多类结构。

Multiclass Classification without Labels via Posterior Simplex Geometry

论文配图:Multiclass Classification without Labels via Posterior Simplex Geometry
图 1 · 摘自论文原文
  • 利用不同混合样本的后验分布几何特性,构建分类器识别混合身份。
  • 在MNIST、CIFAR-10和Galaxy10上成功恢复了真实类别与混合比例。
  • 无需先验信息,适合标签稀缺的多分类发现场景。

在许多分类问题中,实例级标签不可靠或缺失。然而,常可构造弱标注的无标签样本:通过不同筛选条件、来源、人群或实验设置生成的数据集,改变潜在类别比例但不透露具体数值。分类无标签(CWoLa)表明,在二分类情形(K=2)下,仅需区分两个类比例不同的混杂样本,即可恢复最优分类器而无需知道比例。本文将该思想拓展至多分类情形(K>2),其中学习者仅能观测混合身份,既无隐含类别标签,也无类别先验矩阵。我们证明,在多分类混合模型中,贝叶斯最优混合分类器 $g^ ext{⋆}$ 将数据点映射到嵌入于混合后验空间中的 $(K-1)$-单纯形。该单纯形的 $K$ 个顶点由隐含类别通过未知混合矩阵诱导而成。基于此几何结构,我们提出无需先验的训练方法:训练标准分类器以区分混合身份,再通过后处理单纯形拟合或瓶颈架构提取隐含类别结构。在MNIST、CIFAR-10和Galaxy10 DECaLS上的实验显示,仅凭混合身份即可恢复隐含类别及其在混合中的占比。该方法显著缩小了弱监督与全监督性能差距,为标签稀缺领域提供了一个数学严谨且可扩展的多分类发现工具。

原文摘要 · Abstract (English)

In many classification problems, reliable instance-level labels are unavailable. However, it is often possible to construct weakly enriched unlabeled samples: datasets selected by different cuts, sources, populations, or experimental conditions that change latent class proportions without revealing them. Classification without Labels (CWoLa) shows that, in the binary case ($K=2$), a classifier trained to distinguish two impure mixtures with different class proportions can recover an optimal class discriminator without knowing the mixture proportions. We extend this principle to multiclass learning from several unlabeled mixtures ($K>2$), where the learner observes only mixture identity and neither latent class labels nor class-prior matrices. We prove that, for a multiclass mixture model, the Bayes-optimal mixture classifier $g^\star$ maps data points into a $(K-1)$-simplex embedded in mixture-posterior space. The $K$ vertices of this simplex are induced by the latent classes through the unknown mixing matrix. Leveraging this geometry, we propose prior-free procedures that train a standard classifier to distinguish mixture identities and then extract latent class structure using either post-hoc simplex fitting or a bottleneck architecture. Experiments on MNIST, CIFAR-10, and Galaxy10 DECaLS show that mixture identity alone can recover latent classes and their fractions in the mixture. By narrowing the gap between weakly supervised and fully supervised performance, we provide a mathematically grounded, scalable tool for multiclass discovery in label-scarce domains.

无标签学习多分类几何结构聚类发现

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。