arXiv:2412.20644cs.LGstat.ML2024-12ICLR被引 14

一种适用于各种标注预算的高效主动学习方法

Uncertainty Herding: One Active Learning Method for All Label Budgets

  • 提出不确定性覆盖机制,统一低/高预算场景
  • 在多种任务中性能优于或媲美现有最优方法
  • 计算快且无需调参,适合实际应用

多数主动学习研究聚焦于标注预算充足时的表现,但在预算较小时性能显著劣于随机选择;另一些方法专为低预算设计,但随着预算增加表现变差。由于‘低’与‘高’预算的界限因任务而异,这在实践中构成严重问题。本文提出不确定性覆盖(Uncertainty Coverage),一种可统一多种低/高预算目标的优化准则,并设计了无需调参的平滑插值方法,实现低-高预算间的无缝切换。我们称贪婪优化该估计的方法为‘不确定性聚类’(Uncertainty Herding),该方法计算高效,且理论上几乎最优地覆盖分布。在多个主动学习任务上的实验表明,该方法在几乎所有情况下均达到或超越当前最优水平,是目前唯一能稳定适用于低、高预算场景的方法。

原文摘要 · Abstract (English)

Most active learning research has focused on methods which perform well when many labels are available, but can be dramatically worse than random selection when label budgets are small. Other methods have focused on the low-budget regime, but do poorly as label budgets increase. As the line between "low" and "high" budgets varies by problem, this is a serious issue in practice. We propose uncertainty coverage, an objective which generalizes a variety of low- and high-budget objectives, as well as natural, hyperparameter-light methods to smoothly interpolate between low- and high-budget regimes. We call greedy optimization of the estimate Uncertainty Herding; this simple method is computationally fast, and we prove that it nearly optimizes the distribution-level coverage. In experimental validation across a variety of active learning tasks, our proposal matches or beats state-of-the-art performance in essentially all cases; it is the only method of which we are aware that reliably works well in both low- and high-budget settings.

主动学习标注预算无参方法鲁棒性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。