arXiv:2511.17629cs.LG2025-11中稿 · IEEE ISBI 2026被引 1

针对罕见病诊断的极端类别不平衡问题,提出高效增强与过滤方法。

Boundary-Aware Adversarial Filtering for Reliable Diagnosis under Extreme Class Imbalance

  • 先合成少数类样本,再用对抗判别器和边界效用模型筛选
  • 召回率和平均精度均优于SMOTE等主流方法,且校准性最佳
  • 特别适合医疗诊断等漏诊代价高的场景

在极端类别不平衡场景下,如医疗诊断,召回率和校准性均至关重要。本文提出AF-SMOTE,一种数学驱动的增强框架:先合成少数类样本,再通过对抗判别器和边界效用模型进行过滤。在满足决策边界平滑性和类条件密度弱假设下,证明该过滤步骤单调提升F_beta(β≥1)的代理指标,且不增加Brier得分。在MIMIC-IV代理标签预测和典型欺诈检测基准上,AF-SMOTE的召回率和平均精度均高于SMOTE、ADASYN、Borderline-SMOTE、SVM-SMOTE等强基线,并实现最优校准。进一步验证了其在多个非MIMIC-IV数据集上的泛化优势。将AF-SMOTE应用于基于代理标签的医疗数据集,以疾病无关方式展示了其在临床中的实用价值——避免罕见病漏诊带来的严重后果。

原文摘要 · Abstract (English)

We study classification under extreme class imbalance where recall and calibration are both critical, for example in medical diagnosis scenarios. We propose AF-SMOTE, a mathematically motivated augmentation framework that first synthesizes minority points and then filters them by an adversarial discriminator and a boundary utility model. We prove that, under mild assumptions on the decision boundary smoothness and class-conditional densities, our filtering step monotonically improves a surrogate of F_beta (for beta >= 1) while not inflating Brier score. On MIMIC-IV proxy label prediction and canonical fraud detection benchmarks, AF-SMOTE attains higher recall and average precision than strong oversampling baselines (SMOTE, ADASYN, Borderline-SMOTE, SVM-SMOTE), and yields the best calibration. We further validate these gains across multiple additional datasets beyond MIMIC-IV. Our successful application of AF-SMOTE to a healthcare dataset using a proxy label demonstrates in a disease-agnostic way its practical value in clinical situations, where missing true positive cases in rare diseases can have severe consequences.

类别不平衡医疗诊断样本增强模型校准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。