arXiv:2511.19487cs.LGstat.ML2025-11

将随机森林邻近性扩展到任意距离学习场景,支持分类与回归任务。

The Generalized Proximity Forest

  • 提出广义邻近森林模型,通用化随机森林邻近性方法。
  • 在分类和回归任务中均优于随机森林与KNN模型。
  • 可作为元学习框架,提升预训练分类器的缺失值补全能力。

近期研究表明,随机森林(RF)邻近性在异常检测、缺失数据填补和可视化等监督学习任务中具有实用性。然而,其效果依赖于随机森林模型的表现,而该模型并非所有场景下的最优选择。已有工作通过基于距离的邻近森林(PF)模型将随机森林邻近性扩展至时间序列分析。本文提出广义邻近森林模型,将随机森林邻近性推广至所有可进行有监督距离学习的场景。此外,我们还引入了适用于回归任务的PF变体,并提出将广义PF模型作为元学习框架,从而将监督填补能力拓展至任意预训练分类器。实验表明,与随机森林及k-近邻模型相比,广义邻近森林模型展现出独特优势。

原文摘要 · Abstract (English)

Recent work has demonstrated the utility of Random Forest (RF) proximities for various supervised machine learning tasks, including outlier detection, missing data imputation, and visualization. However, the utility of the RF proximities depends upon the success of the RF model, which itself is not the ideal model in all contexts. RF proximities have recently been extended to time series by means of the distance-based Proximity Forest (PF) model, among others, affording time series analysis with the benefits of RF proximities. In this work, we introduce the generalized PF model, thereby extending RF proximities to all contexts in which supervised distance-based machine learning can occur. Additionally, we introduce a variant of the PF model for regression tasks. We also introduce the notion of using the generalized PF model as a meta-learning framework, extending supervised imputation capability to any pre-trained classifier. We experimentally demonstrate the unique advantages of the generalized PF model compared with both the RF model and the $k$-nearest neighbors model.

机器学习邻近性元学习时间序列

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。