arXiv:2603.22750stat.MLcs.LG2026-03被引 1

通过枚举近优模型集,提升决策树主动学习的可解释性与效率。

REALITrees: Rashomon Ensemble Active Learning for Interpretable Trees

  • 构建近优模型集合,以精确刻画假设空间多样性。
  • 在中等噪声环境下,收敛速度比随机集成快30%以上。
  • 适合需要高可解释性的机器学习场景,如医疗决策支持。

主动学习通过选择信息量最大的样本降低标注成本。主流方法查询委员会(QBC)依赖扰动带来的多样性,通过随机特征子集或数据遮蔽诱导模型分歧。尽管这近似了认知不确定性,却牺牲了对合理假设空间的直接表征。本文提出互补方法:拉什蒙德集成主动学习(REAL),通过穷举所有近优模型构成委员会。为解决该集合内的功能冗余,采用基于PAC-Bayesian框架的Gibbs后验对委员会成员按经验风险加权。借助最新算法进展,我们能精确枚举稀疏决策树类别的拉什蒙德集。在合成数据和经典主动学习基线测试中,REAL优于随机集成,在中等噪声环境中利用模型多样性优势,实现更快收敛。

原文摘要 · Abstract (English)

Active learning reduces labeling costs by selecting samples that maximize information gain. A dominant framework, Query-by-Committee (QBC), typically relies on perturbation-based diversity by inducing model disagreement through random feature subsetting or data blinding. While this approximates one notion of epistemic uncertainty, it sacrifices direct characterization of the plausible hypothesis space. We propose the complementary approach: Rashomon Ensembled Active Learning (REAL) which constructs a committee by exhaustively enumerating the Rashomon Set of all near-optimal models. To address functional redundancy within this set, we adopt a PAC-Bayesian framework using a Gibbs posterior to weight committee members by their empirical risk. Leveraging recent algorithmic advances, we exactly enumerate this set for the class of sparse decision trees. Across synthetic and established active learning baselines, REAL outperforms randomized ensembles, particularly in moderately noisy environments where it strategically leverages expanded model multiplicity to achieve faster convergence.

主动学习决策树可解释性模型集成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。