arXiv:2501.10540stat.MLcs.LG2025-01被引 2

针对含缺失值的混合数据,直接估计协方差矩阵,提升精度与效率。

DPERC: Direct Parameter Estimation for Mixed Data

  • 利用类别特征信息增强连续变量的协方差估计
  • 在多个数据集上表现优于现有方法,尤其在缺失数据下更稳定
  • 适合需要高效准确相关性分析的研究者使用

协方差矩阵是主成分分析、相关热力图等众多统计与机器学习应用的基础。然而,数据中的缺失值给其准确估计带来巨大挑战。虽然插补方法可缓解此问题,但常需在计算效率与估计精度间权衡。因此,研究转向直接参数估计,因其兼具高精度与低计算负担。本文提出针对含类别特征的随机缺失数据的直接参数估计方法(DPERC),专为包含缺失连续特征的混合数据设计。该方法通过挖掘类别特征中蕴含的信息,显著提升连续变量协方差矩阵的估计效果。在多种数据集上的全面评估表明,DPERC在性能上具有竞争力,且实验验证其可有效生成高质量相关热力图。

原文摘要 · Abstract (English)

The covariance matrix is a foundation in numerous statistical and machine-learning applications such as Principle Component Analysis, Correlation Heatmap, etc. However, missing values within datasets present a formidable obstacle to accurately estimating this matrix. While imputation methods offer one avenue for addressing this challenge, they often entail a trade-off between computational efficiency and estimation accuracy. Consequently, attention has shifted towards direct parameter estimation, given its precision and reduced computational burden. In this paper, we propose Direct Parameter Estimation for Randomly Missing Data with Categorical Features (DPERC), an efficient approach for direct parameter estimation tailored to mixed data that contains missing values within continuous features. Our method is motivated by leveraging information from categorical features, which can significantly enhance covariance matrix estimation for continuous features. Our approach effectively harnesses the information embedded within mixed data structures. Through comprehensive evaluations of diverse datasets, we demonstrate the competitive performance of DPERC compared to various contemporary techniques. In addition, we also show by experiments that DPERC is a valuable tool for visualizing the correlation heatmap.

协方差估计混合数据缺失值处理相关热图

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。