arXiv:2507.18520stat.MLcs.LG2025-07被引 1

在高维异方差噪声下,可无参数地校正距离计算误差。

Euclidean Distance Deflation Under High-Dimensional Heteroskedastic Noise

  • 提出无超参数的联合估计方法,同时还原噪声强度与真实距离。
  • 理论证明误差以多项式速度收敛,数据量和维度越大越准。
  • 适合单细胞测序等高维噪声数据,提升近邻分析可靠性。

成对欧氏距离计算是众多机器学习与数据分析算法的基础步骤。然而,在真实场景中,这些距离常受异方差噪声(即不同观测点噪声幅度不一)干扰,导致距离被非平凡地放大,扭曲原始数据几何结构。本文研究在缺乏干净数据结构或噪声分布先验知识的前提下,如何准确估计每条观测的噪声幅度并校正成对距离。令人意外的是,我们证明在一般高维设定下,该任务仍可可靠完成,即使噪声水平差异显著。为此,我们提出一种原则性、无超参数的方法,联合估计噪声幅度并修正距离。提供理论保证:在归一化ℓ₁范数下,噪声估计与距离校正误差的概率上界随特征维度与数据集规模增加,以多项式速率趋于零。合成数据实验表明,本方法在挑战性条件下能精确还原距离,显著提升后续距离相关计算的鲁棒性。应用于单细胞RNA测序数据时,所获噪声估计与经典模型一致,实现了准确的最近邻识别,为下游分析奠定基础。

原文摘要 · Abstract (English)

Pairwise Euclidean distance calculation is a fundamental step in many machine learning and data analysis algorithms. In real-world applications, however, these distances are frequently distorted by heteroskedastic noise$\unicode{x2014}$a prevalent form of inhomogeneous corruption characterized by variable noise magnitudes across data observations. Such noise inflates the computed distances in a nontrivial way, leading to misrepresentations of the underlying data geometry. In this work, we address the tasks of estimating the noise magnitudes per observation and correcting the pairwise Euclidean distances under heteroskedastic noise. Perhaps surprisingly, we show that in general high-dimensional settings and without assuming prior knowledge on the clean data structure or noise distribution, both tasks can be performed reliably, even when the noise levels vary considerably. Specifically, we develop a principled, hyperparameter-free approach that jointly estimates the noise magnitudes and corrects the distances. We provide theoretical guarantees for our approach, establishing probabilistic bounds on the estimation errors of both noise magnitudes and distances. These bounds, measured in the normalized $\ell_1$ norm, converge to zero at polynomial rates as both feature dimension and dataset size increase. Experiments on synthetic datasets demonstrate that our method accurately estimates distances in challenging regimes, significantly improving the robustness of subsequent distance-based computations. Notably, when applied to single-cell RNA sequencing data, our method yields noise magnitude estimates consistent with an established prototypical model, enabling accurate nearest neighbor identification that is fundamental to many downstream analyses.

高维数据噪声校正距离度量

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。