arXiv:2507.07771stat.MLcs.LG2025-07

统一弱监督框架,用无标签数据提升N元组学习效果

A Unified Empirical Risk Minimization Framework for Flexible N-Tuples Weak Supervision

  • 基于经验风险最小化,统一处理N元组与无标签数据
  • 在多个基准数据集上,利用无标签数据显著提升泛化性能
  • 适用于多种弱监督场景,特别适合标注成本高的任务

为缓解监督学习中的标注负担,N元组学习作为一种新兴的弱监督方法受到关注。现有方法虽将成对比较扩展至高阶关系并适配多种现实场景,但多依赖特定设计且缺乏统一理论基础。本文提出一种基于经验风险最小化的通用N元组学习框架,系统融合点式无标签数据以提升学习性能。首次在统一的概率模型下描述N元组与点式无标签数据的生成过程,由此推导出一类广义的无偏经验风险估计器,可涵盖多数现有模型。进一步建立泛化误差界提供理论支持。通过四个代表性弱监督场景实例验证框架灵活性,均作为本模型的特例可还原。针对负风险项导致的过拟合问题,引入修正函数调整经验风险。大量实验在基准数据集上验证框架有效性,表明利用点式无标签数据能持续提升各类N元组学习任务的泛化能力。

原文摘要 · Abstract (English)

To alleviate the annotation burden in supervised learning, N-tuples learning has recently emerged as a powerful weakly-supervised method. While existing N-tuples learning approaches extend pairwise learning to higher-order comparisons and accommodate various real-world scenarios, they often rely on task-specific designs and lack a unified theoretical foundation. In this paper, we propose a general N-tuples learning framework based on empirical risk minimization, which systematically integrates pointwise unlabeled data to enhance learning performance. This paper first unifies the data generation processes of N-tuples and pointwise unlabeled data under a shared probabilistic formulation. Based on this unified view, we derive an unbiased empirical risk estimator that generalizes a broad class of existing N-tuples models. We further establish a generalization error bound for theoretical support. To demonstrate the flexibility of the framework, we instantiate it in four representative weakly supervised scenarios, each recoverable as a special case of our general model. Additionally, to address overfitting issues arising from negative risk terms, we adopt correction functions to adjust the empirical risk. Extensive experiments on benchmark datasets validate the effectiveness of the proposed framework and demonstrate that leveraging pointwise unlabeled data consistently improves generalization across various N-tuples learning tasks.

弱监督N元组学习无标签数据风险最小化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。