兼顾标签相关性与数据不平衡,提升多标签主动学习效果
Multi-Label Bayesian Active Learning with Inter-Label Relationships
- 构建动态正负相关矩阵,捕捉标签共现与互斥关系
- 在四个真实数据集上优于现有方法,显著提升标注效率
- 适合处理标签相关性强、数据分布不均的复杂场景
多标签主动学习的核心挑战在于评估不确定数量标签的有用性,同时考虑固有的标签相关性。现有方法或需大量计算资源建模相关性,或未能充分挖掘标签依赖。此外,真实场景中常面临数据分布不平衡带来的固有偏差。本文提出一种新策略,通过逐步更新的正负相关矩阵,捕获已标注样本标签空间中的共现与互斥关系,实现对不确定性的整体评估,而非将标签孤立看待。同时,结合集成伪标签和贝塔评分规则,在保持多样性的同时缓解数据不平衡问题。在四个真实数据集上的大量实验表明,该方法相较多种主流方法,始终展现出更可靠且更优的性能。
原文摘要 · Abstract (English)
The primary challenge of multi-label active learning, differing it from multi-class active learning, lies in assessing the informativeness of an indefinite number of labels while also accounting for the inherited label correlation. Existing studies either require substantial computational resources to leverage correlations or fail to fully explore label dependencies. Additionally, real-world scenarios often require addressing intrinsic biases stemming from imbalanced data distributions. In this paper, we propose a new multi-label active learning strategy to address both challenges. Our method incorporates progressively updated positive and negative correlation matrices to capture co-occurrence and disjoint relationships within the label space of annotated samples, enabling a holistic assessment of uncertainty rather than treating labels as isolated elements. Furthermore, alongside diversity, our model employs ensemble pseudo labeling and beta scoring rules to address data imbalances. Extensive experiments on four realistic datasets demonstrate that our strategy consistently achieves more reliable and superior performance, compared to several established methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。