用合成数据训练决策树,高效实现可解释模型的元学习。
Towards Scalable Meta-Learning of near-optimal Interpretable Models via Synthetic Model Generations
- 通过合成近似最优决策树生成大规模真实数据集。
- 在少量真实数据上表现接近全量真实数据训练效果。
- 适合需要高效、可解释模型的金融与医疗领域应用。
决策树因其可解释性被广泛应用于金融、医疗等高风险领域。本文提出一种高效、可扩展的方法,通过合成生成预训练数据以实现决策树的元学习。该方法可合成近似最优的决策树,构建大规模且真实的训练数据集。结合MetaTree Transformer架构,实验表明该方法在性能上可媲美基于真实世界数据或计算成本高昂的最优决策树的预训练。该策略显著降低计算开销,提升数据生成灵活性,为可解释决策树模型的规模化、高效元学习提供了新路径。
原文摘要 · Abstract (English)
Decision trees are widely used in high-stakes fields like finance and healthcare due to their interpretability. This work introduces an efficient, scalable method for generating synthetic pre-training data to enable meta-learning of decision trees. Our approach samples near-optimal decision trees synthetically, creating large-scale, realistic datasets. Using the MetaTree transformer architecture, we demonstrate that this method achieves performance comparable to pre-training on real-world data or with computationally expensive optimal decision trees. This strategy significantly reduces computational costs, enhances data generation flexibility, and paves the way for scalable and efficient meta-learning of interpretable decision tree models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。