通过筛选独特优质模型提升主动学习准确性与可解释性
Unique Rashomon Sets for Robust Active Learning
- 从近优模型集中挑选差异化的高质模型进行集成
- 在五个数据集上准确率最高提升20%
- 适合追求高精度与可解释性的主动学习研究者
机器学习标注数据的收集成本高昂。主动学习通过选择最具信息量的样本减少标注负担,但初始标注数据有限时,难以区分真正的不确定性与噪声导致的波动。随机森林等集成方法虽能量化不确定性,却会无差别地聚合所有模型,包括表现差和冗余的模型,噪声环境下问题更严重。本文提出UNIQUE Rashomon Ensembled Active Learning(UNREAL),仅对来自近优模型集(Rashomon set)的差异化高质模型进行集成。该策略有助于区分真实不确定性与噪声引起的波动。理论证明UNREAL收敛速度优于传统方法,实验显示其在五个基准数据集上预测准确率最高提升20%,同时增强模型可解释性。
原文摘要 · Abstract (English)
Collecting labeled data for machine learning models is often expensive and time-consuming. Active learning addresses this challenge by selectively labeling the most informative observations, but when initial labeled data is limited, it becomes difficult to distinguish genuinely informative points from those appearing uncertain primarily due to noise. Ensemble methods like random forests are a powerful approach to quantifying this uncertainty but do so by aggregating all models indiscriminately. This includes poor performing models and redundant models, a problem that worsens in the presence of noisy data. We introduce UNique Rashomon Ensembled Active Learning (UNREAL), which selectively ensembles only distinct models from the Rashomon set, which is the set of nearly optimal models. Restricting ensemble membership to high-performing models with different explanations helps distinguish genuine uncertainty from noise-induced variation. We show that UNREAL achieves faster theoretical convergence rates than traditional active learning approaches and demonstrates empirical improvements of up to 20% in predictive accuracy across five benchmark datasets, while simultaneously enhancing model interpretability.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。