在线时间序列中,让模型自动决定何时自判、何时求助专家。
Learning-to-Defer in Non-Stationary Time Series via Switching State-Space Models
- 用切换状态空间模型建模数据变化,融合内部预测与外部专家。
- 真实数据上仅需少于2%的轮次调用专家,性能优于基线。
- 适合在线学习、数据漂移场景,如金融、气象预测。
学习性延迟(L2D)将每个决策交给系统自身预测器或外部专家。流式时间序列场景打破了传统离线L2D假设:数据非平稳,专家可用性随时间变化,且内部预测器在线训练。我们提出L2D-SLDS,一种基于因子分解切换线性高斯状态空间模型的一阶段在线L2D框架,涵盖所有潜在残差:离散模式、共享全局因子和各专家特有状态。始终可得的内部残差通过共享因子持续更新对未查询专家的信念,学习者感知查询评分则平衡即时成本与潜在状态信息增益及一步学习改进。我们证明了对时变学习-延迟比较器的奥拉克不等式,将遗憾分解为查询奖励预算、SLDS预测误差项$ \mathcal{E}_{\mathrm{SLDS}}$以及内部学习者的区间动态遗憾。在合成数据、墨尔本、耶纳及24位专家的德里基准上,L2D-SLDS表现媲美或超越上下文与非平稳贝叶斯基线,同时在真实数据上仅需少于2%的轮次调用专家。
原文摘要 · Abstract (English)
Learning-to-defer (L2D) routes each decision to a system's own predictor or to an external expert. Streaming time-series settings break the offline-L2D assumptions: the data are non-stationary, expert availability shifts over time, and the internal predictor is trained online. We propose L2D-SLDS, a one-stage online L2D framework based on a factorized switching linear-Gaussian state-space model over all potential residuals: a discrete regime, a shared global factor, and per-expert idiosyncratic states. The always-observed internal residual continuously updates beliefs about every unqueried expert through the shared factor, and a learner-aware query score balances immediate cost against latent-state information gain and one-step learner improvement. We prove an oracle inequality against a time-varying learn-and-defer comparator, decomposing regret into a query-bonus budget, an SLDS predictive-cost-error term~$\mathcal{E}_{\mathrm{SLDS}}$, and the internal learner's interval dynamic regret. On synthetic, Melbourne, Jena, and 24-expert Delhi benchmarks, L2D-SLDS is competitive with or improves on contextual- and non-stationary-bandit baselines while deferring on ${<}2\%$ of real-data rounds.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。