提出用费雪-罗伊距离量化分布漂移,揭示学习者行为对数据分布的影响机制。
Learning under Distributional Drift: Prequential Reproducibility as an Intrinsic Statistical Resource
- 引入漂移预算 $C_T$,用费雪-罗伊距离衡量数据分布的累积几何运动
- 证明性能预测误差上界为 $T^{-1/2} + C_T/T$,且该依赖关系紧致
- 适用于闭环学习、反馈扰动等场景,适合研究自适应系统稳定性的人
在分布漂移环境下进行统计学习仍缺乏清晰刻画,尤其在学习过程改变数据生成规律的闭环设置中。本文引入一个内在漂移预算 $C_T$,通过费雪-罗伊距离量化沿学习者-环境轨迹的数据分布累积信息几何运动。该预算区分了外部环境变化与学习者行为引发的策略敏感反馈。由此得到预序可复现性的速率表征:当利用实际序列上的表现预测下一分布的一步前瞻表现时,漂移影响由平均运动率 $C_T/T$ 决定,而非累计漂移本身。我们证明了漂移-反馈误差界为 $T^{-1/2} + C_T/T$,并针对一个典型正则子类建立了匹配的紧下界。因此,$C_T/T$ 的依赖关系在常数意义下是紧的——既足够用于上界控制,又在困难正则子类中不可避免。进一步,我们建立了一个信息论不可区分性结果:一阶 $C/T$ 效应无法仅从实际性能流中识别。最后,我们表明固定监控通道会压缩可观测的费雪运动,并在实验中(包括一个误设的真实数据反馈场景)验证,合理选择通道可在原始数据生成规律不可知时保留关键风险信号。理论统一了外生漂移、自适应数据分析和表现反馈作为同一轨迹上的费雪-罗伊运动的不同来源。
原文摘要 · Abstract (English)
Statistical learning under distributional drift remains poorly characterized, especially in closed-loop settings where learning alters the data-generating law. We introduce an intrinsic drift budget $C_T$ that quantifies cumulative information-geometric motion of the data distribution along the realized learner-environment trajectory, measured in Fisher-Rao distance. The budget separates exogenous environmental change from policy-sensitive feedback induced by the learner's actions. This gives a rate-based characterization of prequential reproducibility: when performance on the realized stream is used to predict one-step-ahead performance under the next distribution, the drift contribution enters through the average motion rate $C_T/T$, not through cumulative drift alone. We prove a drift-feedback bound of order $T^{-1/2}+C_T/T$, up to controlled second-order remainder terms, and establish a matching sharpness lower bound for the same prequential reproducibility gap on a canonical regular subclass. Thus the dependence on the average Fisher-Rao motion rate is tight up to constants: $C_T/T$ is sufficient for upper control and unavoidable on regular hard subclasses. We further prove an information-theoretic indistinguishability result showing that order-$C/T$ effects on the one-step-ahead target need not be identifiable from the realized performance stream alone. Finally, we show that fixed monitoring channels induce contracted observable Fisher motion, and experiments, including a misspecified real-data feedback setting, indicate that appropriately chosen channels can retain risk-relevant drift signal when the intrinsic data-generating law is unavailable. The resulting theory treats exogenous drift, adaptive data analysis, and performative feedback as different sources of Fisher-Rao motion along the same learner-environment trajectory.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。