arXiv:2606.07630cs.LGcs.AI2026-06

用大模型先验提升小模型在噪声与不均衡数据下的标注效率

Active Learning with Foundation Model Priors: Efficient Learning under Class Imbalance

论文配图:Active Learning with Foundation Model Priors: Efficient Learning under Class Imbalance
图 1 · 摘自论文原文
  • 结合大模型先验,让小模型与大模型协同判断样本价值
  • 标注量减少超50%,且在噪声数据中表现更鲁棒
  • 适合处理真实世界中标签混乱、类别不均的数据场景

图像和文本领域的现实数据集常存在类别分布偏斜和标注噪声,二者共同导致模型性能下降,尤其影响少数类表现。现有解决方案中,主动学习通过有选择地查询最具信息量且平衡的样本进行标注,提供了一种高效范式。本文提出一种创新的主动学习框架,有效缓解类别不平衡问题,并筛选最有价值的样本进行标注。利用大模型先验,算法实现大模型与小模型之间的不平衡感知联合决策,以应对跨图像与文本领域中的噪声与不平衡标签。我们首次系统研究了在标签噪声与类别不平衡双重挑战下,主动学习在图像与文本领域的应用。在多个不均衡数据集上的大量实验表明,该方法相较最优基线主动学习方法,可实现超过50%的标注量节省,同时保持模型性能并具备对标签噪声的鲁棒性。

原文摘要 · Abstract (English)

Real-world datasets across image and text domains are often characterized by skewed class distributions and noisy annotations, which jointly degrade model performance, particularly on minority classes. Among existing solutions, active learning offers an effective and efficient paradigm by selectively querying the most informative and balanced samples for annotation. We propose an innovative active learning framework that mitigates class imbalance and selects the most informative samples to annotate. Leveraging foundation model priors, our algorithm enables imbalance-aware co-decisions between foundation model and small model to tackle noisy and imbalanced labels across various domains. We introduce the first study to systematically explore active learning under the dual challenges of label noise and class imbalance across image and text domains. Extensive experiments on imbalanced datasets demonstrate that our method achieves substantial annotation savings-over 50% compared to the best active learning baseline-while preserving performance and robustness to label noise.

主动学习大模型先验类别不平衡标注效率

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。