arXiv:2604.02019cs.LG2026-04

通过特征加权提升回归主动学习的样本选择精度

Feature Weighting Improves Pool-Based Sequential Active Learning for Regression

论文配图:Feature Weighting Improves Pool-Based Sequential Active Learning for Regression
图 1 · 摘自论文原文
  • 用岭回归系数为特征赋权,改进样本间距离计算
  • 在五种现有方法上均显著提升回归模型性能
  • 适合需要高效标注的回归任务研究者

基于池的序列主动学习用于回归(ALR)从大量未标注样本中逐次挑选少量样本进行标注,以在给定标注预算下构建更精确的回归模型。代表性与多样性是关键考量因素,涉及样本间的距离计算。然而,以往方法未考虑不同特征在距离计算中的重要性,导致距离不准确,进而影响样本选择效果。本文提出四种特征加权单任务ALR方法和三种特征加权多任务ALR方法,利用少量已标注样本训练的岭回归系数对相应特征加权,以改进样本间距离计算。大量实验表明,这一直观且易实现的改进几乎总能提升五种现有ALR方法在单任务和多任务回归问题上的表现。该特征加权策略也可轻松扩展至流式主动学习及分类算法。

原文摘要 · Abstract (English)

Pool-based sequential active learning for regression (ALR) optimally selects a small number of samples sequentially from a large pool of unlabeled samples to label, so that a more accurate regression model can be constructed under a given labeling budget. Representativeness and diversity, which involve computing the distances among different samples, are important considerations in ALR. However, previous ALR approaches do not incorporate the importance of different features in inter-sample distance computation, resulting in inaccurate distances and hence sub-optimal sample selection. This paper proposes four feature weighted single-task ALR approaches and three feature weighted multi-task ALR approaches, where the ridge regression coefficients trained from a small amount of previously labeled samples are used to weight the corresponding features in inter-sample distance computation. Extensive experiments showed that this intuitive and easy-to-implement enhancement almost always improves the performance of five existing ALR approaches, in both single-task and multi-task regression problems. The feature weighting strategy may also be easily extended to stream-based ALR, and classification algorithms.

主动学习回归特征加权优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。