提出新方法,用总体标签比例更准地训练分类模型。
Optimal Learning from Label Proportions with General Loss Functions
- 设计低方差去偏算法,适配多种损失函数。
- 理论证明样本复杂度显著降低,提升学习效率。
- 在多个数据集上表现优于传统方法,适合实际应用。
受在线广告问题启发,本文研究从标签比例学习(LLP)任务。提出一种新颖且通用的低方差去偏方法,可有效利用聚合标签信息,显著提升现有 LLP 方法的性能。该方法具有高度灵活性,能自然适配二分类和多分类场景中的多种实用损失函数。通过与标准技术结合,我们为一大类实际相关的损失函数提供了更优的样本复杂度保证。同时,在多个基准数据集上进行了实证验证,结果表明所提方法在实践中显著优于标准基线。
原文摘要 · Abstract (English)
Motivated by problems in online advertising, we address the task of Learning from Label Proportions (LLP). We introduce a novel and versatile low-variance debiasing methodology to learn from aggregate label information, significantly advancing the state of the art in LLP. Our debiasing approach exhibits remarkable flexibility, seamlessly accommodating a broad spectrum of practically relevant loss functions across both binary and multi-class classification settings. By carefully combining our estimators with standard techniques, we improve sample complexity guarantees for a large class of losses of practical relevance. We also empirically validate the efficacy of our proposed approach across a diverse array of benchmark datasets, demonstrating compelling empirical advantages over standard baselines.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。