arXiv:2602.04042cs.LGstat.ME2026-02

提出可统一处理连续与分类变量的条件密度估计新方法

Partition Tree: Conditional Density Estimation over General Outcome Spaces

  • 基于数据自适应划分建模分段常数密度,直接优化负对数似然
  • 在多个数据集上优于CART类树,性能媲美最先进概率树方法
  • 适合需要精准概率预测的回归与分类任务

我们提出Partition Tree,一种面向通用结果空间的新型树形框架,用于条件密度估计,可在统一框架内处理连续与类别变量。该方法将条件分布建模为数据自适应划分上的分段常数密度,并通过直接最小化条件负对数似然来学习树结构。这提供了一种无需对目标分布做参数假设的可扩展非参数替代方案。我们进一步引入Partition Forest,通过平均条件密度获得袋装扩展。实验表明,其在概率预测上优于CART类树,在性能上与最先进的概率树方法及Random Forests相当。

原文摘要 · Abstract (English)

We propose Partition Tree, a novel tree-based framework for conditional density estimation over general outcome spaces that supports both continuous and categorical variables within a unified formulation. Our approach models conditional distributions as piecewise-constant densities on data-adaptive partitions and learns trees by directly minimizing conditional negative log-likelihood. This yields a scalable, nonparametric alternative to existing probabilistic trees that does not make parametric assumptions about the target distribution. We further introduce Partition Forest, a bagging extension obtained by averaging conditional densities. Empirically, we demonstrate improved probabilistic prediction over CART-style trees and competitive performance compared to state-of-the-art probabilistic tree methods and Random Forests.

密度估计概率树非参数

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。