arXiv:2608.08424stat.MLcs.CV2026-08被引 1

提出ARC方法,让变点定位在有限样本下保持可靠覆盖且对分布变化更鲁棒。

ARC: Augmented-Rank Conformalization for Changepoint Localization --- Finite-Sample Validity and Distribution-Robust Efficiency

  • 基于段内秩构造新评分,仅依赖数据顺序信息
  • 在真实数据上定位准确率达3-5个候选点,空集可标识拟合失败
  • 对单调变换不变,适合分布漂移场景的稳健分析

置信集方法可将任意评分转化为具有有限样本覆盖的变点置信区域,覆盖性普遍成立但效率不保证。最优评分是似然比,实际评分需估计密度比,但在重尾、偏斜和分布漂移下置信集长度恶化且无长度保证。本文提出ARC(增强秩共形化),一类仅通过段内秩定义的评分:秩CUSUM位置与尺度通道、固定组合及轻量神经网络(合成训练后冻结)。所有ARC评分在任意冻结权重配置下(包括随机初始化与错误训练)均保持有限样本覆盖。核心结果为效率传递定理:整个ARC置信集几乎必然在严格递增边际变换下不变,因此其长度分布仅依赖于数据的秩结构,一旦认证即在单调轨道上完全保持,而插值评分的长度随表达方式改变而膨胀。不同秩结构下长度仍可报告。经典秩检验理论表明,ARC以可控成本逼近最优不变评分。模拟验证所有评分均实现名义覆盖,即使网络被破坏;在单调变换下保持一致,而插值评分会膨胀;插值置信集平滑退化至空集。在井孔数据基准上,ARC可将标注变点定位至3-5个候选,空集提示拟合不良。两个边界明确:序列相关破坏精确性,趋势型备择假设超出分段交换模型范围。

原文摘要 · Abstract (English)

Conformal changepoint localization turns any score into a confidence set for the changepoint with finite-sample coverage. Coverage is universal; efficiency is not. The oracle score is a likelihood ratio, so practical scores estimate density ratios, and set length deteriorates under heavy tails, skewness, and distribution shift, where no length guarantee applies. We propose ARC (Augmented-Rank Conformalization), a family of scores depending on the data only through within-segment ranks: rank-CUSUM location and scale channels, their fixed combinations, and a lightweight neural score frozen after synthetic training. Every ARC score inherits finite-sample coverage for every frozen weight configuration, including random initialization and mistraining. The main result is an efficiency transfer theorem: the entire ARC confidence set is almost surely invariant under strictly increasing marginal transforms, so the set length distribution depends on the data pair only through its rank structure, and lengths certified once hold verbatim across its monotone orbit, whereas a plug-in score's length changes with every re-expression. Across different rank structures lengths do change, and are reported as such. Classical rank-test theory positions ARC as targeting the optimal invariant score at bounded cost. Simulations confirm nominal coverage for all scores, including sabotaged networks, identical sets under monotone transforms where plug-in scores inflate, and smooth degradation where plug-in sets become vacuous; on the well-log benchmark ARC localizes annotated shifts to three to five candidates and flags misfit by an empty set. Two boundaries are stated rather than hidden: serial dependence destroys exactness, and trend-type alternatives lie outside the piecewise-exchangeable model.

变点检测共形推断稳健统计秩方法

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。