arXiv:2607.09250stat.MLcs.LG2026-07

高维模型中,样本影响力分布可精确刻画,助力识别关键数据点。

Influence Diagnostics in High-dimensional M-estimation: Precise Asymptotics

  • 基于高维极限下凸M估计的理论分析,推导样本影响力分布极限
  • 影响力大的样本更靠近决策边界,与主动学习策略一致
  • 适用于高维统计建模中的异常点检测与数据筛选

训练样本对统计模型的影响通常通过删除该样本后模型性能的变化来衡量。在低维大样本情形(n→∞, d=O(1))下,这种留一法影响的统计特性已被充分理解;但在高维情形(n≈d)下,单个样本的影响会与其他所有样本产生复杂依赖关系。针对高斯设计下的凸M估计,在高维极限下,我们证明了训练集上影响力分布收敛到一个可精确刻画的极限测度。基于此结果,我们发现具有较大影响力的样本倾向于位于决策边界附近,这与主动学习中常见的数据选择启发式相吻合。

原文摘要 · Abstract (English)

The impact of a given training point on a statistical model is classically measured through its leave-one-out influence, which quantifies the effect of its removal from the training set on the model accuracy. While the statistics of leave-one-out influences are well understood in the low-dimensional, large sample limit $n\to \infty, d=O(1)$, they become more intricate in high dimensions, as the influence of a given sample develops non-trivial dependencies on all other training samples. For convex M-estimation under Gaussian design, in the high-dimensional limit $n\asymp d$, we show that the distribution of the influences across the training set converges to a limiting measure which we sharply characterize. Building on these results, we provide evidence that influential samples tend to lie close to the decision boundary, thereby making contact with a standard data selection heuristic in active learning.

高维统计影响力诊断M估计主动学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。