arXiv:2510.09246cs.LG2025-10

用主成分分析构建数据子空间,预测缺失值。

A PCA-based Data Prediction Method

  • 基于主成分的子空间距离计算缺失数据
  • 使用欧氏距离求解,数学与机器学习结合
  • 适合处理结构化数据缺失问题

数据科学中常面临缺失数据取值选择的问题。本文提出一种融合传统数学与机器学习元素的新方法,用于缺失数据的预测(填补)。该方法基于表示现有数据和候选集的平移线性子空间之间的距离概念。现有数据集由其前几个主成分张成的子空间表示。给出了欧氏距离情况下的求解方案。

原文摘要 · Abstract (English)

The problem of choosing appropriate values for missing data is often encountered in the data science. We describe a novel method containing both traditional mathematics and machine learning elements for prediction (imputation) of missing data. This method is based on the notion of distance between shifted linear subspaces representing the existing data and candidate sets. The existing data set is represented by the subspace spanned by its first principal components. Solutions for the case of the Euclidean metric are given.

数据填补主成分分析子空间距离

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。