检测高维数据中隐藏的三阶分布变化,不依赖密度估计。
High-Dimensional Nonparametric Change-Point Detection via Low-Rank Degree-Three Density Projection
- 用三阶勒让德张量编码密度投影,避免直接估计密度。
- 在200维下仍可准确定位变化点,小跳跃时定位精度达O(κ⁻²)。
- 适合处理高维非参数变化检测,尤其关注偏度和非对称交互。
分布变化可能在均值和协方差上不可见,却体现在偏度、非对称交互或其他三阶结构中。我们提出一种非参数变化点检测方法,保留密度至多三阶的所有系数,同时避免直接密度估计。对于取值于$[-1,1]^d$的观测,构造对称三阶勒让德特征张量$H_3(X)\in\Sym^3(\R^{d+1})$,使得$A(f)=\E_fH_3(X)$为密度三阶投影的精确等距编码:$\|A(f)-A(g)\|_{\F}=\|P_3(f-g)\|_{L^2}$。相反,固定张量收缩是三阶多项式混沌,具有$ψ_{2/3}$尾部。两项均具备三阶张量标度特性,与简单随机张量的尖锐集中结果中的幂次一致。在坐标正交特化下,界优化至$\sqrt{\log d}$,并支持数百维度的前缀和实现。我们推导出理论人口帐篷形状与定位边界,引入带填充局部重中心化的种子最短区间算法,并通过归纳法证明精确恢复:零递归段保持非激活,每个未检测到的变化保留平衡隔离区间,最短活跃种子在重中心化前恰好包含一个变化点。双向交叉拟合标量精修在小跳跃情形下达到$O_{\Pp}(κ^{-2})$定位精度,匹配纯立方族上的Le Cam下界(其二阶投影跳跃恰好为零)。在$d\in\{20,50,100,200\}$及三变化点的$d=100$序列上重现实验,展示了预期高维场景,且未实际生成$(d+1)^3$张量。
原文摘要 · Abstract (English)
Distributional changes can be invisible to means and covariances yet appear in skewness, asymmetric interactions, or other third-order structure. We develop a nonparametric change-point method that retains every degree-at-most-three coefficient of a density while avoiding direct density estimation. For observations in $[-1,1]^d$, we construct a symmetric order-three Legendre feature tensor $H_3(X)\in\Sym^3(\R^{d+1})$ such that $A(f)=\E_fH_3(X)$ is an exact isometric encoding of the degree-three density projection: $\|A(f)-A(g)\|_{\F}=\|P_3(f-g)\|_{L^2}$. Instead, fixed tensor contractions are degree-three polynomial chaoses with $ψ_{2/3}$ tails. The two terms have the characteristic order-three tensor scaling and match the powers in sharp concentration results for simple random tensors. For a coordinate-orthogonal specialization, the bound improves to $\sqrt{\log d}$ and enables a prefix-sum implementation in hundreds of dimensions. We derive the exact population tent shape and localization margin, introduce a seeded shortest-interval algorithm with a padded local recentering step, and prove exact recovery by induction: null recursive segments remain inactive, every undetected change retains a balanced isolating interval, and the shortest active seed contains exactly one change before recentering. A two-way cross-fitted scalar refinement attains $O_{\Pp}(κ^{-2})$ localization in the small-jump regime, matching a Le Cam lower bound on a pure cubic family whose degree-two projection jump is exactly zero. Reproducible experiments at $d\in\{20,50,100,200\}$ and a three-change $d=100$ sequence demonstrate the intended high-dimensional regime without materializing a $(d+1)^3$ tensor.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。