arXiv:2502.11324stat.MLcs.LG2025-02被引 1

高维数据下小样本均值估计的鲁棒性实证研究

Robust High-Dimensional Mean Estimation With Low Data Size, an Empirical Study

  • 对比多种均值估计方法在低样本高维下的表现
  • 发现现有理论最优算法在小样本时性能显著下降
  • 为小样本场景提供实用算法选择参考

稳健统计旨在处理部分数据被任意破坏时的数据表征。最核心的统计量是均值,近年来针对高维受污染数据的均值估计已有大量理论进展,提出了若干近似最优误差的高效算法。然而这些算法普遍依赖于与维度相关的较大样本量要求。本文对多种均值估计技术进行了广泛实验,聚焦于高维设置下数据量不足的情况,评估其在小样本条件下的实际表现。

原文摘要 · Abstract (English)

Robust statistics aims to compute quantities to represent data where a fraction of it may be arbitrarily corrupted. The most essential statistic is the mean, and in recent years, there has been a flurry of theoretical advancement for efficiently estimating the mean in high dimensions on corrupted data. While several algorithms have been proposed that achieve near-optimal error, they all rely on large data size requirements as a function of dimension. In this paper, we perform an extensive experimentation over various mean estimation techniques where data size might not meet this requirement due to the high-dimensional setting.

均值估计高维数据鲁棒性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。