用任务驱动的动态表示提升混乱数据池的主动学习效果
Active Learning with Task-Driven Representations for Messy Pools
- 在主动学习中动态更新任务相关的数据表示
- 相比固定无监督表示,准确率显著提升
- 适合处理标签不齐、质量参差的数据池
主动学习在杂乱、未清理的数据池中具有巨大潜力,其中样本对目标任务的相关性差异较大。然而,现有方法通常依赖固定的无监督表示,仅通过调整选择策略来应对。我们发现这种设置会削弱其在复杂数据池中的表现,因为固定表示难以捕捉与任务相关的关键信息。为此,我们提出在主动学习过程中周期性地利用已收集标签更新任务驱动的表示。设计了两种具体策略:一种是直接学习半监督表示,另一种是微调初始无监督表示。实验表明,这两种方法在性能上均显著优于使用无监督或预训练表示。
原文摘要 · Abstract (English)
Active learning has the potential to be especially useful for messy, uncurated pools where datapoints vary in relevance to the target task. However, state-of-the-art approaches to this problem currently rely on using fixed, unsupervised representations of the pool, focusing on modifying the acquisition function instead. We show that this model setup can undermine their effectiveness at dealing with messy pools, as such representations can fail to capture important information relevant to the task. To address this, we propose using task-driven representations that are periodically updated during the active learning process using the previously collected labels. We introduce two specific strategies for learning these representations, one based on directly learning semi-supervised representations and the other based on supervised fine-tuning of an initial unsupervised representation. We find that both significantly improve empirical performance over using unsupervised or pretrained representations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。