arXiv:2505.21441stat.MLcs.AI2025-05NeurIPS被引 3

用随机森林实现数据压缩与重构,可还原原始特征并用于可视化分析

Autoencoding Random Forests

  • 基于非参数统计和图谱理论,学习数据的低维嵌入表示
  • 通过约束优化与近邻回归实现精确与近似解码,重建输入数据
  • 适用于监督/无监督任务,适合处理表格、图像及基因组数据

我们提出一种针对随机森林的系统性自编码方法。该方法基于非参数统计与谱图理论,学习数据关系的低维嵌入表示。通过约束优化、分裂重标记与最近邻回归,提供解码问题的精确与近似解法。这些方法能有效逆向压缩流程,建立从嵌入空间到输入空间的映射,利用集成中各树的分裂规则进行重构。在常见正则性假设下,解码器具有普遍一致性。该方法适用于监督或无监督模型,可揭示条件或联合分布。实验展示了其在可视化、压缩、聚类与去噪等任务中的强大能力,在表格、图像及基因组数据上均表现出易用性与有效性。

原文摘要 · Abstract (English)

We propose a principled method for autoencoding with random forests. Our strategy builds on foundational results from nonparametric statistics and spectral graph theory to learn a low-dimensional embedding of the model that optimally represents relationships in the data. We provide exact and approximate solutions to the decoding problem via constrained optimization, split relabeling, and nearest neighbors regression. These methods effectively invert the compression pipeline, establishing a map from the embedding space back to the input space using splits learned by the ensemble's constituent trees. The resulting decoders are universally consistent under common regularity assumptions. The procedure works with supervised or unsupervised models, providing a window into conditional or joint distributions. We demonstrate various applications of this autoencoder, including powerful new tools for visualization, compression, clustering, and denoising. Experiments illustrate the ease and utility of our method in a wide range of settings, including tabular, image, and genomic data.

随机森林自编码数据压缩可视化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。