为提升模型稳定性,研究主动学习中可复现性的代价
The Cost of Replicability in Active Learning
- 用随机阈值法设计可复现的主动学习算法
- 在可实现与泛化设定下,仍能显著减少标注需求
- 适合关注模型稳定性的机器学习研究者
主动学习通过有选择地查询未标记数据的标签,降低机器学习算法对标注数据的需求。确保可复现性(即算法在不同运行中产生一致结果)对模型可靠性至关重要,但常导致样本复杂度上升。本文研究了两种经典的基于分歧的主动学习方法(CAL 和 A^2)在可复现性约束下的代价。通过引入随机阈值技术,提出了两类可复现的主动学习算法:适用于有限假设类可实现学习的情形,以及适用于泛化设置的情形。理论分析表明,尽管强制可复现性会增加标签复杂度,但 CAL 与 A^2 仍能在该约束下实现显著的标签节省。研究结果为平衡主动学习中的效率与稳定性提供了新见解。
原文摘要 · Abstract (English)
Active learning aims to reduce the number of labeled data points required by machine learning algorithms by selectively querying labels from initially unlabeled data. Ensuring replicability, where an algorithm produces consistent outcomes across different runs, is essential for the reliability of machine learning models but often increases sample complexity. This paper investigates the cost of replicability in active learning using two classical disagreement-based methods: the CAL and A^2 algorithms. Leveraging randomized thresholding techniques, we propose two replicable active learning algorithms: one for realizable learning of finite hypothesis classes and another for the agnostic setting. Our theoretical analysis shows that while enforcing replicability increases label complexity, CAL and A^2 still achieve substantial label savings under this constraint. These findings provide insights into balancing efficiency and stability in active learning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。