arXiv:2608.13922cs.LGstat.ML2026-08

通过低秩投影检测高维分布突变,能发现均值方法忽略的依赖关系变化。

High-dimensional nonparametric changepoint detection via low-rank degree-two density projection

  • 用矩阵均值替代密度估计,保留二阶密度信息
  • 在高维数据中准确识别出均值不变但依赖结构变化的突变点
  • 适合检测隐藏在复杂依赖中的突变,尤其对非参数场景有效

在预变化和后变化密度均无参数形式的情况下,高维分布突变检测极具挑战。本文提出一种基于表示的方法,通过构造对称特征矩阵 $H_2(X) o bR^{(d+1) imes(d+1)}$,将密度的二阶正交投影等距编码为矩阵期望 $M(f) = bE_f H_2(X)$,从而避免直接密度估计。对矩阵进行秩-$r$截断后,扫描矩阵CUSUM统计量,利用投影跳跃的低秩特性而非坐标稀疏性。所提出的 exttt{LRD} 估计器具有帐篷形总体目标函数,并具备非渐近算子范数分析,其主导随机项规模为 $ ilde O( oot r d rac{ d}{ m n})$。对于多突变情形,提出种子最小过阈值法,通过归纳证明可精确恢复所有突变点,且保持未检测突变的隔离区间。交叉拟合标量精修方法在一个折上学习变化方向,另一折进行定位,达到 $ ilde O_{bP}(κ^{-2})$ 的误差界;匹配的Le Cam下界表明该结果在对数因子内最优。针对几何β-混合情形,借助依赖矩阵Bernstein不等式实现扩展。实验涵盖最高200维、$d=100$的三突变序列及128维人类活动基准数据集,验证方法计算可行,且能精准检测均值不变但依赖结构改变的突变,而此类变化对均值CUSUM方法完全不可见。

原文摘要 · Abstract (English)

Detecting distributional changes in high dimension is difficult when neither the pre-change nor post-change density is parametrically specified. We introduce a representation-based approach that retains all degree-at-most-two density information while replacing density estimation by matrix mean estimation. For observations in $[-1,1]^d$, a symmetric feature matrix $H_2(X)\in\R^{(d+1)\times(d+1)}$ is constructed so that $M(f)=\E_f H_2(X)$ is an isometric encoding of the degree-two orthogonal projection of the density. We scan matrix CUSUMs after rank-$r$ truncation, exploiting the low rank of the projected jump rather than sparsity of individual coordinates. The resulting \LRD{} estimator has a tent-shaped population objective and a nonasymptotic operator-norm analysis whose leading stochastic term scales as $\sqrt{rd\log(nd)}$. For multiple changes, we give a seeded narrowest-over-threshold procedure and prove exact recovery by an induction that preserves an isolating interval for every undetected change. A cross-fitted scalar refinement learns the changing low-rank direction on one fold and localizes on the other, attaining $\widetilde O_{\Pp}(κ^{-2})$ error; a matching Le Cam lower bound shows optimality up to logarithms. A geometrically $β$-mixing extension follows from a dependent matrix Bernstein inequality. Experiments with ambient dimension up to $200$, a three-change $d=100$ sequence, and a $128$-feature human-activity benchmark show that the method remains computationally practical and accurately detects pure dependence changes that are invisible to mean CUSUMs.

突变检测高维统计低秩模型非参数

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。