解决异常样本稀少且分散的双重不平衡问题,提升小样本异常检测精度。
Detecting Scarce and Sparse Anomalous: Solving Dual Imbalance in Multi-Instance Learning
- 将多实例学习重构为细粒度正未标记学习,从微观层面平衡数据分布。
- 在真实与合成数据集上均显著优于现有方法,尤其在极稀疏异常场景下表现突出。
- 适合处理医疗影像、工业质检等异常样本极少的实际应用场景。
在真实应用中,检测具有极度稀疏异常的样本极具挑战性,因其与正常样本高度相似,易被误判。同时,异常样本本身数量稀少,导致多实例学习面临宏观与微观双重不平衡问题。为此,本文提出将该问题重新建模为细粒度正未标记学习,通过微观层面的平衡机制实现无偏处理。基于严谨理论基础,我们设计了新型框架BFGPU。大量实验在合成与真实数据集上验证了其有效性,显著提升了极稀疏异常检测性能。
原文摘要 · Abstract (English)
In real-world applications, it is highly challenging to detect anomalous samples with extremely sparse anomalies, as they are highly similar to and thus easily confused with normal samples. Moreover, the number of anomalous samples is inherently scarce. This results in a dual imbalance Multi-Instance Learning (MIL) problem, manifesting at both the macro and micro levels. To address this "needle-in-a-haystack problem", we find that MIL problem can be reformulated as a fine-grained PU learning problem. This allows us to address the imbalance issue in an unbiased manner using micro-level balancing mechanisms. To this end, we propose a novel framework, Balanced Fine-Grained Positive-Unlabeled (BFGPU)-based on rigorous theoretical foundations. Extensive experiments on both synthetic and real-world datasets demonstrate the effectiveness of BFGPU.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。