针对开放集标注中的噪声问题,提出基于狄利克雷分布的分级选例方法。
Dirichlet-Based Coarse-to-Fine Example Selection For Open-Set Annotation
- 用狄利克雷分布建模不确定性,打破软最大值带来的平移不变性
- 通过双头模型差异识别难样本,结合不确定性和差异性分两阶段选例
- 在多种开放度数据集上表现优于现有方法,适合开放集场景下的主动学习
主动学习在从无标签数据中选择最有价值样本方面取得了显著成效。然而,在涉及开放集噪声的真实场景中,其性能会显著下降,这一问题被称为开放集标注(OSA)。本文将性能下降归因于基于软最大值的平移不变性导致的不可靠预测,为此提出基于狄利克雷分布的粗到细选例(DCFS)策略。该方法引入基于单纯形的证据深度学习(EDL),通过同时考虑证据数据和分布不确定性,打破平移不变性,区分已知与未知类别。此外,通过两个分类器头生成的模型差异识别困难的已知类样本,并分别放大和抑制未知与已知类的差异。最终,将差异与不确定性结合,形成两级选例策略,从已知类别中选择最具信息量的样本。在多种开放度比率数据集上的大量实验表明,DCFS达到当前最优性能。
原文摘要 · Abstract (English)
Active learning (AL) has achieved great success by selecting the most valuable examples from unlabeled data. However, they usually deteriorate in real scenarios where open-set noise gets involved, which is studied as open-set annotation (OSA). In this paper, we owe the deterioration to the unreliable predictions arising from softmax-based translation invariance and propose a Dirichlet-based Coarse-to-Fine Example Selection (DCFS) strategy accordingly. Our method introduces simplex-based evidential deep learning (EDL) to break translation invariance and distinguish known and unknown classes by considering evidence-based data and distribution uncertainty simultaneously. Furthermore, hard known-class examples are identified by model discrepancy generated from two classifier heads, where we amplify and alleviate the model discrepancy respectively for unknown and known classes. Finally, we combine the discrepancy with uncertainties to form a two-stage strategy, selecting the most informative examples from known classes. Extensive experiments on various openness ratio datasets demonstrate that DCFS achieves state-of-art performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。