解决预训练数据不平衡下的开放世界模型适应问题
OASIS: Open-world Adaptive Self-supervised and Imbalanced-aware System
- 基于对比学习的预训练提升少数类识别能力
- 生成可靠伪标签,显著提升开放世界场景准确率
- 自适应激活机制降低计算开销,适合实时系统
机器学习向动态环境扩展时面临开放世界问题,如标签分布变化、特征分布偏移及未知类别出现。现有后训练方法在初始预训练数据存在类别不平衡时表现受限,难以泛化到少数类。本文提出一种新方法,在不平衡数据预训练基础上仍能有效应对开放世界挑战。通过对比学习增强预训练,显著提升对少数类的分类性能;设计后训练机制生成可靠伪标签,增强模型鲁棒性;引入选择性激活策略优化训练过程,减少无效计算。大量实验表明,该方法在多种开放世界场景中均显著优于当前最优适配技术,在准确率和效率上均有明显提升。
原文摘要 · Abstract (English)
The expansion of machine learning into dynamic environments presents challenges in handling open-world problems where label shift, covariate shift, and unknown classes emerge. Post-training methods have been explored to address these challenges, adapting models to newly emerging data. However, these methods struggle when the initial pre-training is performed on class-imbalanced datasets, limiting generalization to minority classes. To address this, we propose a method that effectively handles open-world problems even when pre-training is conducted on imbalanced data. Our contrastive-based pre-training approach enhances classification performance, particularly for underrepresented classes. Our post-training mechanism generates reliable pseudo-labels, improving model robustness against open-world problems. We also introduce selective activation criteria to optimize the post-training process, reducing unnecessary computation. Extensive experiments demonstrate that our method significantly outperforms state-of-the-art adaptation techniques in both accuracy and efficiency across diverse open-world scenarios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。