arXiv:2505.05355cs.LG2025-05ICML被引 4

提出近似最优的标签比例学习样本复杂度,提升小样本下的模型精度。

Nearly Optimal Sample Complexity for Learning with Label Proportions

  • 基于经验风险最小化与随机梯度下降改进算法,结合方差缩减技术。
  • 样本复杂度随袋大小增长速率接近理论最优,优于现有方法。
  • 适合数据标注不完整但需高精度的小样本场景,如医疗分析。

我们研究标签比例学习(LLP),一种部分信息设定:训练样本被分组为袋子,仅知每袋中各类别标签的总和。尽管观测不完整,目标仍是实现个体样本上的小损失。在平方损失下,我们给出了LLP的样本复杂度结果,表明其近乎最优。算法上,采用经精心设计的ERM与随机梯度下降变体,并引入专门的方差减少技术。理论上,我们的结果在样本复杂度对袋大小的依赖关系上显著改进了现有文献;实验上,在多个数据集上验证了算法性能,相比近期基线,在更少样本下实现了更高准确率。

原文摘要 · Abstract (English)

We investigate Learning from Label Proportions (LLP), a partial information setting where examples in a training set are grouped into bags, and only aggregate label values in each bag are available. Despite the partial observability, the goal is still to achieve small regret at the level of individual examples. We give results on the sample complexity of LLP under square loss, showing that our sample complexity is essentially optimal. From an algorithmic viewpoint, we rely on carefully designed variants of Empirical Risk Minimization, and Stochastic Gradient Descent algorithms, combined with ad hoc variance reduction techniques. On one hand, our theoretical results improve in important ways on the existing literature on LLP, specifically in the way the sample complexity depends on the bag size. On the other hand, we validate our algorithmic solutions on several datasets, demonstrating improved empirical performance (better accuracy for less samples) against recent baselines.

标签比例样本效率机器学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。