arXiv:2409.09078stat.MLcs.LG2024-09被引 1

为主动学习设计可验证的查询策略,提升模型泛化能力

Bounds on the Generalization Error in Active Learning

  • 通过积分概率度量融合信息性与代表性,设计更优查询策略
  • 证明正则化能确保泛化误差上界成立,适用于多种假设类
  • 提供理论工具评估查询算法质量,指导实际应用设计

本文通过推导泛化误差的一族上界,建立了主动学习中的经验风险最小化原理。结果表明,结合信息性与代表性查询策略可获得更优的查询算法,其中代表性通过积分概率度量评估。我们系统地将不同损失函数和假设类对应的主动学习场景与其相应的上界相联系。研究显示,用于约束各类假设类复杂度的正则化技术是保证上界有效性的重要充分条件。本工作实现了主动学习中查询算法的理论化构建与实证质量评估。

原文摘要 · Abstract (English)

We establish empirical risk minimization principles for active learning by deriving a family of upper bounds on the generalization error. Aligning with empirical observations, the bounds suggest that superior query algorithms can be obtained by combining both informativeness and representativeness query strategies, where the latter is assessed using integral probability metrics. To facilitate the use of these bounds in application, we systematically link diverse active learning scenarios, characterized by their loss functions and hypothesis classes to their corresponding upper bounds. Our results show that regularization techniques used to constraint the complexity of various hypothesis classes are sufficient conditions to ensure the validity of the bounds. The present work enables principled construction and empirical quality-evaluation of query algorithms in active learning.

主动学习泛化误差正则化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。