提出统一加权框架,解决小样本与数据不平衡问题。
FOSSIL: Regret-minimizing weighting for robust learning under imbalance and small data
- 基于后悔最小化设计统一加权公式,整合难例学习与增强惩罚
- 在真实与合成数据上均优于传统方法,无需修改网络结构
- 适合罕见病影像、基因组等小样本高不平衡场景
小样本与数据不平衡广泛存在于罕见病影像、基因组学和灾害响应等领域,标注样本稀缺且盲目增强常引入伪影。现有方法如过采样、焦点损失或元加权仅解决部分问题,仍易失效或复杂。本文提出FOSSIL(基于样本敏感重要性学习的灵活优化),一个统一的加权框架,将类别不平衡校正、难度感知课程学习、增强惩罚和预热动态整合为单一可解释公式。相比以往启发式方法,该框架提供基于后悔的理论保障,在合成与真实数据集上持续优于ERM、课程学习及元加权基线,且无需架构改动。
原文摘要 · Abstract (English)
Imbalanced and small data regimes are pervasive in domains such as rare disease imaging, genomics, and disaster response, where labeled samples are scarce and naive augmentation often introduces artifacts. Existing solutions such as oversampling, focal loss, or meta-weighting address isolated aspects of this challenge but remain fragile or complex. We introduce FOSSIL (Flexible Optimization via Sample Sensitive Importance Learning), a unified weighting framework that seamlessly integrates class imbalance correction, difficulty-aware curricula, augmentation penalties, and warmup dynamics into a single interpretable formula. Unlike prior heuristics, the proposed framework provides regret-based theoretical guarantees and achieves consistent empirical gains over ERM, curriculum, and meta-weighting baselines on synthetic and real-world datasets, while requiring no architectural changes.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。