通过设计非均匀的错误标签生成方式,显著提升多类别弱监督学习效果。
Embracing Biased Transition Matrices for Complementary-Label Learning with Many Classes

- 主动引入标签偏差,限制错误标签仅来自部分类别
- 在CIFAR-100和TinyImageNet-200上实现超7倍准确率提升
- 为真实场景中多类别弱监督学习提供新路径
互补标签学习(CLL)是一种弱监督范式,其中样本被标注为不属于的类别。尽管已有十年研究,现有方法在10类分类任务中仍具竞争力,但在大规模标签空间下难以扩展,这主要源于传统方法假设标签生成过程均匀分布,导致多类别设置下学习信号严重稀释。本文表明,通过有意识地设计非均匀(有偏)的生成过程,将互补标签限制在特定子集内,可有效突破这一长期瓶颈。受此启发,我们提出了一种系统性框架BICL(Bias-Induced Constrained Labeling),覆盖数据收集到训练全过程,充分利用该偏差机制。BICL在CIFAR-100和TinyImageNet-200上实现了超过七倍于传统方法的准确率提升,为使CLL在真实应用中适用于多类别场景开辟了新方向。
原文摘要 · Abstract (English)
Complementary-label learning (CLL) is a weakly supervised paradigm where instances are labeled with classes they do not belong to. Despite a decade of research, CLL methods remain competitive mainly on 10-class classification, with scaling to large label spaces continuing to be an enduring bottleneck. This limitation stems from the common assumption of uniform label generation in traditional methods, which fatally dilutes the learning signal in many-class settings. In this paper, we demonstrate that this long-standing barrier can be overcome by deliberately designing a biased (non-uniform) generation process that restricts complementary labels to a subset of classes. This finding motivates us to propose Bias-Induced Constrained Labeling (BICL), a principled framework spanning data collection to training that leverages this bias. BICL enables effective learning on CIFAR-100 and TinyImageNet-200, achieving more than sevenfold accuracy improvements over traditional methods. Our findings establish a new trajectory for making CLL feasible for many classes in real-world applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。