提出统一方法解决不精确标注下的多标签分类问题。
Rethinking Consistent Multi-Label Classification Under Inexact Supervision
- 设计一阶与二阶风险估计器,无需依赖标签生成假设。
- 理论证明对常用评估指标具一致性,并给出误差收敛速率。
- 在真实与合成数据上优于现有方法,适合弱监督场景。
部分多标签学习和互补多标签学习是两种流行的弱监督多标签分类范式,旨在降低精确标注多标签数据的高成本。在部分多标签学习中,每个实例被标注一个候选标签集,其中仅部分标签相关;在互补多标签学习中,每个实例被标注其不属于的类别。现有的一致性方法要么需要准确估计候选或互补标签的生成过程,要么假设均匀分布以避免估计问题,但这些条件在真实场景中通常难以满足。本文提出无需依赖上述条件的统一一致性方法。具体而言,基于一阶与二阶策略设计两种风险估计器。理论上,证明了该方法相对于两种广泛使用的多标签分类评估指标的一致性,并推导出所提风险估计器的估计误差收敛率。实验上,大量在真实与合成数据集上的结果验证了所提方法相较于现有最优方法的有效性。
原文摘要 · Abstract (English)
Partial multi-label learning and complementary multi-label learning are two popular weakly supervised multi-label classification paradigms that aim to alleviate the high annotation costs of collecting precisely annotated multi-label data. In partial multi-label learning, each instance is annotated with a candidate label set, among which only some labels are relevant; in complementary multi-label learning, each instance is annotated with complementary labels indicating the classes to which the instance does not belong. Existing consistent approaches for the two paradigms either require accurate estimation of the generation process of candidate or complementary labels or assume a uniform distribution to eliminate the estimation problem. However, both conditions are usually difficult to satisfy in real-world scenarios. In this paper, we propose consistent approaches that do not rely on the aforementioned conditions to handle both problems in a unified way. Specifically, we propose two risk estimators based on first- and second-order strategies. Theoretically, we prove consistency w.r.t. two widely used multi-label classification evaluation metrics and derive convergence rates for the estimation errors of the proposed risk estimators. Empirically, extensive experimental results on both real-world and synthetic datasets validate the effectiveness of our proposed approaches against state-of-the-art methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。