用额外监督信息提升少量标注数据下的分类准确率
Learning from Hard Labels with Additional Supervision on Non-Hard-Labeled Classes
- 将硬标签与附加监督融合生成软标签,构建概率分布模型
- 附加监督中非硬标签类的信息才是提升性能的关键
- 理论揭示了监督方向与步长的互补作用,适合小样本场景
在观测成本高或数据稀缺的场景下,丰富每条样本的标签信息对构建高精度分类模型至关重要。此时,除硬标签外,常可获取诸如硬标签置信度等附加监督信息。这引发两个核心问题:哪些附加监督本质上有益?它们如何促进泛化性能提升?为此,我们提出一个理论框架,将硬标签与附加监督均视为概率分布,并通过仿射组合构建软标签。理论分析表明,附加监督的核心价值不在于所赋硬标签的置信度,而在于非硬标签类上的分布信息。此外,我们证明附加监督与混合系数在软标签优化中发挥互补作用:前者决定从硬标签分布向真实分布调整的方向,后者控制沿该方向的步长。通过泛化误差分析,我们理论刻画了附加监督及其混合系数对误差界收敛速度与渐近值的影响。实验验证了基于该理论设计的附加监督,即使简单使用也能显著提升分类准确率。
原文摘要 · Abstract (English)
In scenarios where training data is limited due to observation costs or data scarcity, enriching the label information associated with each instance becomes crucial for building high-accuracy classification models. In such contexts, it is often feasible to obtain not only hard labels but also {\it additional supervision}, such as the confidences for the hard labels. This setting naturally raises fundamental questions: {\it What kinds of additional supervision are intrinsically beneficial?} And {\it how do they contribute to improved generalization performance?} To address these questions, we propose a theoretical framework that treats both hard labels and additional supervision as probability distributions, and constructs soft labels through their affine combination. Our theoretical analysis reveals that the essential component of additional supervision is not the confidence score of the assigned hard label, but rather the information of the distribution over the non-hard-labeled classes. Moreover, we demonstrate that the additional supervision and the mixing coefficient contribute to the refinement of soft labels in complementary roles. Intuitively, in the probability simplex, the additional supervision determines the direction in which the deterministic distribution representing the hard label should be adjusted toward the true label distribution, while the mixing coefficient controls the step size along that direction. Through generalization error analysis, we theoretically characterize how the additional supervision and its mixing coefficient affect both the convergence rate and asymptotic value of the error bound. Finally, we experimentally demonstrate that, based on our theory, designing additional supervision can lead to improved classification accuracy, even when utilized in a simple manner.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。