定位数据异常源头,让统计检验结果可解释。
Localizing Global Discrepancies: Marginal Contributions and Contextual Anomaly Detection
- 通过随机上下文计算每条数据的边际或条件贡献,实现异常定位。
- 在大型强子对撞机数据集上,局部估计器与真实值相关性达0.9993。
- 揭示上下文信息可提供独立特征外的结构化判别能力。
全局拟合优度和差异统计量可判断样本是否偏离参考分布,但无法定位具体驱动异常的观测值。本文提出一种定位框架,通过在随机统计上下文中为每条观测分配其条件或边际贡献,将重抽样诊断、数据估值与投影理论及事件级异常检测相连接。对于对称统计量,固定大小替换等价于中心化条件定位;对于U统计量,新增得分等于首个Hoeffding/Hájek贡献;对于平滑分布泛函,其在主导阶上与影响函数相关;对于无偏已知背景的MMD,精确退化为MMD见证函数。该视角还带来更高效的估计器:匹配上下文相减可消除与观测无关的波动,而配对式MMD中的事件包含项可作为简单局部器。在LHC奥运会异常检测基准上,配对估计器以1/(Rm²)的预测速率收敛至直接经验MMD见证函数,当批次大小m=1000、批次数R=5×10⁶时,相关性达0.9993,且AUC基本一致。我们进一步探讨了上下文是否蕴含事件自身特征之外的信息。在共享潜在变量的模拟模型中,单事件信号与背景分布构造相同,导致孤立事件的AUC=0.5。区分信息仅存在于由共享潜在参数引起的跨事件依赖中;全体样本能恢复此信息,而独立潜在控制组则不能。这区分了上下文的双重作用:高效定位全局差异,以及在备择假设含共享结构时提供真正额外分类信息。
原文摘要 · Abstract (English)
Global goodness-of-fit and discrepancy statistics can establish that a sample departs from a reference distribution without identifying which observations drive the departure. We develop a framework for this localization problem by assigning to each observation its conditional or marginal contribution across random statistical contexts. This connects resampling diagnostics and data valuation to projection theory and event-level anomaly detection. For symmetric statistics, fixed-size replacement is exactly equivalent to centered conditional localization. For U-statistics, the addition score equals the first Hoeffding/Hájek contribution; for smooth distributional functionals it is related at leading order to the influence function; and for unbiased known-background MMD it reduces exactly to the MMD witness. This viewpoint also yields more efficient estimators. Matched-context subtraction removes fluctuations unrelated to the observation, while for pairwise MMD the event-containing terms give a simple localizer. On the LHC Olympics anomaly-detection benchmark, the pair estimator converges to the direct empirical MMD witness with the predicted 1/(Rm^2) scaling, where m is batch size and R the number of batches. At m=1000 and R=5x106 it reaches correlation 0.9993 with essentially identical AUC. We also ask when context contains information beyond an event's own features. In a shared-latent toy model, the full single-event signal and background distributions are identical by construction, forcing isolated-event AUC=0.5. Discriminating information survives only in cross-event dependence induced by the shared latent parameter; the ensemble recovers this information, whereas an independent-latent control does not. This separates two roles of context: efficient localization of a global discrepancy and genuinely additional class information when the alternative contains shared structure.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。