arXiv:2501.15790cs.LGstat.ML2025-01被引 1

提出可为每个样本提供安全保证的合成数据生成方法,兼顾性能与可靠性。

Certified Interpolation Oversampling: Per-Instance Safety Guarantees for Imbalanced Learning

  • 基于安全引导的三阶段插值框架,确保合成样本的安全性
  • 在45个数据集上表现媲美SMOTE,梯度提升下排名第二
  • 首次实现每样本的可认证安全性,适合高风险场景使用

合成少数类样本通常以提升分类性能为目标,本文另设一个目标:生成具有明确安全属性的样本,且该属性对每个实例均通过构造保证而非假设。提出认证插值安全采样(CISO),一种三阶段插值框架。安全引导分布决定合成位置,局部性与清空加权分布选择插值对象,q-高斯分布密度确定实例在连线上的放置位置。该框架提供三项保障:由足够安全的锚点生成的合成实例与多数类保持已认证距离;符号温度参数可严格单调地调节边界或内部探索模式;选择权重恒为正,避免退化情况。认证与优异预测性能并行达成。在评估前注册的协议下,覆盖45个数据集、4种分类器和11种对比方法,CISO在精度-召回曲线下面积上与SMOTE统计等效,在梯度提升中位列第二,11,460次交叉验证全部成功完成。参数扫描进一步揭示了连续的保真-安全权衡,而竞争方法仅能占据孤立点。

原文摘要 · Abstract (English)

Synthetic minority oversampling is typically designed and evaluated against a predictive objective, generating samples that improve downstream classification. This paper pursues a second objective by generating samples that carry a stated safety property, established for each instance by construction rather than assumed. We introduce Certified Interpolation Safe Oversampling (CISO), a three-phase interpolation framework built for this objective. A safety-guided distribution selects where synthesis occurs, a locality-and clearance-weighted distribution selects with whom each anchor interpolates, and a q-Gaussian placement density determines how far along the resulting segment each instance is placed. The framework provides three guarantees. Each synthetic instance seeded by a sufficiently safe anchor carries a certified distance from the majority class; a signed temperature parameter provably and monotonically shifts synthesis between boundary-seeking and interior-seeking regimes; and selection weights are strictly positive by construction, so no degenerate case arises. Certification is obtained alongside competitive predictive performance rather than in place of it. Under a protocol preregistered before evaluation, across 45 datasets, four classifiers, and eleven competing methods, CISO is statistically equivalent to SMOTE on precision-recall AUC, ranks second of eleven under gradient boosting, and completes every one of 11,460 fold-level evaluations without failure. A parameter sweep further reveals a continuous fidelity-safety trade-off that competing methods occupy only as isolated points.

数据增强安全学习不平衡数据认证生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。