提出首个深度部分标签学习基准,解决评估不公与数据不足问题。
Realistic Evaluation of Deep Partial-Label Learning Algorithms
- 构建首个系统性对比深度部分标签算法的基准PLENCH
- 发现早期算法常被低估,部分可超越复杂新模型
- 创建真实标注的人工智能标注图像数据集PLCIFAR10
部分标签学习(PLL)是一种弱监督学习任务,每个样本关联多个候选标签,但仅有一个为真标签。近年来许多深度PLL算法被提出以提升性能,但我们发现一些早期算法常被低估,甚至优于设计复杂的后续算法。本文从实证角度深入分析PLL,揭示三个关键但长期被忽视的问题:第一,模型选择对PLL非平凡,却从未系统研究;第二,实验设置极不一致,难以公平评估算法效果;第三,缺乏适配现代网络架构的真实世界图像数据集。基于此,我们提出PLENCH——首个部分标签学习基准,首次系统研究PLL的模型选择问题,并提出具备理论保证的新标准。同时构建了人工标注的部分标签图像数据集PLCIFAR10,源自Amazon Mechanical Turk。研究者可基于PLENCH快速、便捷地开展全面且公平的评估,验证新算法的有效性。我们希望未来能推动PLL算法的标准化、公平化和实用性评估。
原文摘要 · Abstract (English)
Partial-label learning (PLL) is a weakly supervised learning problem in which each example is associated with multiple candidate labels and only one is the true label. In recent years, many deep PLL algorithms have been developed to improve model performance. However, we find that some early developed algorithms are often underestimated and can outperform many later algorithms with complicated designs. In this paper, we delve into the empirical perspective of PLL and identify several critical but previously overlooked issues. First, model selection for PLL is non-trivial, but has never been systematically studied. Second, the experimental settings are highly inconsistent, making it difficult to evaluate the effectiveness of the algorithms. Third, there is a lack of real-world image datasets that can be compatible with modern network architectures. Based on these findings, we propose PLENCH, the first Partial-Label learning bENCHmark to systematically compare state-of-the-art deep PLL algorithms. We investigate the model selection problem for PLL for the first time, and propose novel model selection criteria with theoretical guarantees. We also create Partial-Label CIFAR-10 (PLCIFAR10), an image dataset of human-annotated partial labels collected from Amazon Mechanical Turk, to provide a testbed for evaluating the performance of PLL algorithms in more realistic scenarios. Researchers can quickly and conveniently perform a comprehensive and fair evaluation and verify the effectiveness of newly developed algorithms based on PLENCH. We hope that PLENCH will facilitate standardized, fair, and practical evaluation of PLL algorithms in the future.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。