提出可随时验证的量化预测审计方法,能识别不同信息水平下的预测偏差。
Bet on Features: Anytime-Valid and Feature-Aware Auditing of Conditional Quantile Forecasters

- 基于博弈论设计无分布假设的连续审计框架,支持非独立同分布数据。
- 在真实数据中发现主流时间序列模型对多个特征存在显著预测偏差。
- 结果按特征层级展示,适合关注预测可靠性与公平性的从业者。
黑箱条件分位数预测广泛应用于供应链管理等具有非对称成本的序列决策场景。部署后需持续监控因数据流漂移和模式变化导致的标准固定周期回测失效的问题。此外,现有回测方法未考虑校准性依赖于审计者信息水平:粗粒度信息下看似校准的预测,在更丰富信息下可能失准。本文提出一种无需分布假设、基于博弈论的连续审计框架,适用于非独立同分布损失,可有效检测由审计者可用特征所指定的可预测替代假设。我们形式化了不同特征集下的条件分位数校准概念,揭示审计者信息越粗,测试难度越高。针对线性依赖于特征的上下文投注,推导出有限时间内的检测保证,且无需i.i.d.假设。所得证据过程可在特征层面解释,量化细粒度的“特征感知”校准偏差证据。在模拟与真实数据上验证方法,发现主流时间序列模型Chronos-2在多个相关特征上严重失准。
原文摘要 · Abstract (English)
Black-box conditional quantile forecasts are widely used for sequential decisions under asymmetric costs, such as inventory planning in supply chain management. Once deployed, such forecasters must be monitored continuously as data streams drift and regimes change; this invalidates standard, fixed-horizon backtests for calibration. Further, existing backtests do not take into account that the notion of calibration is, in fact, information-dependent: forecasts can look calibrated to an auditor with coarse information while being miscalibrated to an auditor with richer information. We develop a distribution-free and game-theoretic testing framework for continuously auditing black-box conditional quantile forecasters with non-i.i.d. losses, such that the resulting evidence process is powerful against predictably chosen alternatives specified by the features available to the auditor. We first formalize notions of conditional quantile calibration when different sets of features are available to the auditor, establishing that the coarseness of the auditor's information set determines the hardness of the testing problem. We then identify the sets of alternatives for which the auditor can achieve power, and focusing on contextual bets linear in the features, we derive finite-time detection guarantees for such alternatives, all without an i.i.d. assumption. The resulting evidence processes are interpretable at the feature level, as they quantify fine-grained, "feature-aware" evidence for miscalibration. We empirically validate these methods on simulated and real data, finding that a popular time series forecaster (Chronos-2) is highly miscalibrated w.r.t. multiple relevant features.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。