arXiv:2606.00512cs.LGcs.IT2026-06

用噪声代理变量提升低标签数据下的回归性能

Semi-Supervised Learning with Noisy Proxy Covariates: Generalization Bounds and Distribution Regression

  • 先提取所有代理变量的核特征,再用岭回归拟合标签数据
  • 当代理变量扰动可控且无标签数据充足时,可实现快速标签样本率
  • 特别适合标签极少但有大量预训练特征的场景

在现代机器学习流程中,大量预训练表示常作为噪声代理变量,而任务标签却十分稀缺。本文研究在此设定下的半监督回归问题,提出一种简单两阶段估计器:首先从所有代理变量中学习核特征,然后在有标签数据上拟合岭回归模型。我们推导了有限样本泛化界,表明当代理变量扰动受控且无标签代理变量足够多时,可恢复快速的标签样本率。我们还证明分布回归是该方法的直接特例,在袋大小足够大时具有类似保证。实验显示,该方法在低标签情形下持续优于监督与半监督基线。

原文摘要 · Abstract (English)

In many modern machine learning pipelines, abundant pretrained representations serve as noisy proxy covariates, while task-specific labels remain scarce. We study semi-supervised regression in this setting, and propose a simple two stage estimator that learns kernel eigenfeatures from all proxy covariates and fits a ridge predictor on labeled data. We derive finite sample bounds showing that fast labeled sample rates are recovered when proxy perturbation is controlled and unlabeled proxy covariates are sufficiently abundant. We also show that distribution regression is a direct special case, with analogous guarantees when the finite bag size is large enough. Experiments show consistent gains over supervised and semi-supervised baselines, especially in low label regimes.

半监督学习回归代理变量泛化界

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。