arXiv:2603.15321cs.LG2026-03被引 1

提出跨模型类的高效近似最优模型集合,支持多视角建模选择。

CASHomon Sets: Efficient Rashomon Sets Across Multiple Model Classes and their Hyperparameters

  • 基于主动学习估计隐式阈值的水平集,实现多模型类联合搜索
  • 在真实与合成数据上识别出高质量替代模型,性能优于多种基线方法
  • 揭示单一模型类解释的局限性,适合需多视角决策的研究者

Rashomon 集合是同一模型类别中性能接近参考模型的一组模型,它们的存在表明存在多种表现良好但解释可能不同的模型,可用于匹配领域知识或用户偏好。然而,当前高效构建方法仅适用于少数模型类别。实际机器学习常需在多个模型类别中搜索,且最佳类别事先未知。因此,本文研究算法选择与超参数优化(CASH)场景下的 Rashomon 集合,称之为 CASHomon 集合。提出 TruVaRImp——一种基于模型的主动学习算法,用于隐式阈值的水平集估计,并提供收敛性保证。在合成和真实数据集上,TruVaRImp 能可靠识别 CASHomon 集合成员,性能匹配或超越朴素采样、贝叶斯优化、经典及隐式水平集估计方法等基线。对预测多重性和特征重要性变异性的分析质疑了仅依赖单一模型类进行数据分析的常见做法。

原文摘要 · Abstract (English)

Rashomon sets are model sets within one model class that perform nearly as well as a reference model from the same model class. They reveal the existence of alternative well-performing models, which may support different interpretations. This enables selecting models that match domain knowledge, hidden constraints, or user preferences. However, efficient construction methods currently exist for only a few model classes. Applied machine learning usually searches many model classes, and the best class is unknown beforehand. We therefore study Rashomon sets in the combined algorithm selection and hyperparameter optimization (CASH) setting and call them CASHomon sets. We propose TruVaRImp, a model-based active learning algorithm for level set estimation with an implicit threshold, and provide convergence guarantees. On synthetic and real-world datasets, TruVaRImp reliably identifies CASHomon sets members and matches or outperforms naive sampling, Bayesian optimization, classical and implicit level set estimation methods, and other baselines. Our analyses of predictive multiplicity and feature-importance variability across model classes question the common practice of interpreting data through a single model class.

模型选择不确定性主动学习可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。