arXiv:2508.00586cs.LG2025-08

主动学习在低数据场景中效果有限,但配合数据增强与半监督学习可提升性能。

The Role of Active Learning in Modern Machine Learning

  • 将主动学习与数据增强、半监督学习结合,提升模型性能。
  • 主动学习仅比随机采样提升1-4%,而增强与半监督方法可提升达60%。
  • 主动学习适合作为最后优化步骤,而非解决标签缺失的首选方案。

尽管主动学习(AL)被广泛研究,但在其学术文献之外很少应用。我们认为原因在于其高计算成本及在少量标注数据下带来的收益较小。本文研究了应对低数据场景的不同方法:数据增强(DA)、半监督学习(SSL)和主动学习(AL)。结果表明,主动学习是解决低数据问题中最不高效的方法,仅比随机采样提升1-4%;而结合随机采样时,数据增强与半监督学习可实现高达60%的性能提升。然而,当主动学习与强数据增强和半监督技术结合时,仍能带来进一步改进。据此,我们提出应将主动学习视为在应用合适的数据增强与半监督方法后,用于榨取数据最后性能潜力的最终工具。

原文摘要 · Abstract (English)

Even though Active Learning (AL) is widely studied, it is rarely applied in contexts outside its own scientific literature. We posit that the reason for this is AL's high computational cost coupled with the comparatively small lifts it is typically able to generate in scenarios with few labeled points. In this work we study the impact of different methods to combat this low data scenario, namely data augmentation (DA), semi-supervised learning (SSL) and AL. We find that AL is by far the least efficient method of solving the low data problem, generating a lift of only 1-4\% over random sampling, while DA and SSL methods can generate up to 60\% lift in combination with random sampling. However, when AL is combined with strong DA and SSL techniques, it surprisingly is still able to provide improvements. Based on these results, we frame AL not as a method to combat missing labels, but as the final building block to squeeze the last bits of performance out of data after appropriate DA and SSL methods as been applied.

主动学习数据增强半监督学习低数据

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。