arXiv:2502.13257cs.LG2025-02NeurIPS被引 7

用随机森林自编码器实现可扩展的监督可视化,提升准确性和解释性。

Random Forest Autoencoders for Guided Representation Learning

论文配图:Random Forest Autoencoders for Guided Representation Learning
图 1 · 摘自论文原文
  • 结合自编码器与随机森林,构建可外推的监督可视化框架
  • 在准确性和可解释性上优于现有方法,尤其适合小样本场景
  • 对超参数不敏感,通用性强,适配多种降维方法

大量研究已发展出稳健的无监督数据可视化方法。然而,专家标签引导的有监督可视化仍鲜受关注,因多数方法侧重分类而非可视化。近期,基于随机森林与信息几何的扩散流形学习方法RF-PHATE在有监督可视化上取得显著进展。但其缺乏显式映射函数,限制了扩展性与对未见数据的应用,难以应对大规模数据集和标签稀缺场景。为此,我们提出随机森林自编码器(RF-AE),一种基于神经网络的框架,用于外推核函数,融合自编码器的灵活性、随机森林的监督学习优势以及RF-PHATE所捕捉的几何结构。RF-AE实现了高效的外推有监督可视化,在准确性和可解释性上均优于现有方法,包括标准核扩展版本的RF-PHATE。此外,该方法对超参数选择鲁棒,可推广至任意基于核的降维方法。

原文摘要 · Abstract (English)

Extensive research has produced robust methods for unsupervised data visualization. Yet supervised visualization$\unicode{x2013}$where expert labels guide representations$\unicode{x2013}$remains underexplored, as most supervised approaches prioritize classification over visualization. Recently, RF-PHATE, a diffusion-based manifold learning method leveraging random forests and information geometry, marked significant progress in supervised visualization. However, its lack of an explicit mapping function limits scalability and its application to unseen data, posing challenges for large datasets and label-scarce scenarios. To overcome these limitations, we introduce Random Forest Autoencoders (RF-AE), a neural network-based framework for out-of-sample kernel extension that combines the flexibility of autoencoders with the supervised learning strengths of random forests and the geometry captured by RF-PHATE. RF-AE enables efficient out-of-sample supervised visualization and outperforms existing methods, including RF-PHATE's standard kernel extension, in both accuracy and interpretability. Additionally, RF-AE is robust to the choice of hyperparameters and generalizes to any kernel-based dimensionality reduction method.

监督可视化随机森林自编码器降维

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。