arXiv:2606.07910cs.LG2026-06中稿 · the NYRL 2025 Work…

用上下文动态选最优主动学习策略,提升标注效率

CAAL: Contextual Bandits based Online Hand-Craft Active Learning Strategy Selection

论文配图:CAAL: Contextual Bandits based Online Hand-Craft Active Learning Strategy Selection
图 1 · 摘自论文原文
  • 将主动学习策略视为可选‘臂’,基于外部上下文预测奖励来选策略
  • 在多个公开数据集上优于现有自适应策略,且批次大小不影响效果
  • 支持领域知识注入,适合需要高效标注的工业级场景

主动学习算法面临未标注数据分布不确定的挑战,难以选出最优的手动设计策略。为此,我们提出上下文自适应主动学习(CAAL)。在CAAL中,每个“臂”代表一个手动设计的策略。不同于仅依赖已标注数据反馈的现有框架,我们通过外部上下文信息进行奖励预测,动态选择用于标注数据批次的策略。该通用框架支持结合领域知识设计更有效的奖励和上下文候选。实验表明,在使用我们设计的奖励与上下文时,CAAL在多个公开数据集上均优于现有基线自适应策略,且结果在不同迭代批次大小下保持一致。

原文摘要 · Abstract (English)

The challenge with active learning algorithms is the uncertainty of the statistical distribution of unlabeled data, making it difficult to choose the best hand-crafted strategy. To address this, we introduced Contextual Adaptive Active Learning (CAAL). In CAAL, each "arm" represents a hand-crafted strategy. Unlike existing frameworks that select strategies based only on feedback from labeled data, we dynamically choose strategies for labeling batches of data using reward prediction with external context information. This general framework allows for customization with domain knowledge to design more effective rewards and context candidates. In addition, we experimentally show that CAAL outperforms the existing baseline adaptive strategy on public datasets using our reward and context design. Our results are consistent regardless of batch size in each iteration.

主动学习强化学习上下文感知标注优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。