arXiv:2607.22774cs.LG2026-07

提出可同时适应成本与时间的高效停时决策方法,解决传统方法需重复训练的问题。

CC-AOS: Cost- and Horizon-Conditioned Amortized Backward Induction for Finite-Horizon Optimal Stopping

论文配图:CC-AOS: Cost- and Horizon-Conditioned Amortized Backward Induction for Finite-Horizon Optimal Stopping
图 1 · 摘自论文原文
  • 用条件化神经网络联合学习多成本多时域下的继续值函数
  • 在福特发动机噪声数据集上,6个未见组合平均降本15.75%
  • 适合需要动态调整采样成本和预测周期的实时时序决策场景

有限时域最优停时是早期时序分类的核心问题,系统需在每个序列前缀判断继续观测的期望收益是否超过采样成本。现有数据驱动的反向归纳方法通常对每个成本-时域组合单独求解,导致改变运行条件需重复优化并维护多套模型,难以实现连续成本调节和多时域部署。本文提出CC-AOS(成本与时域条件化摊销最优停时),一种针对连续成本和多时域组合的结构化摊销求解器。通过联合摊销拟合反向归纳,学习一个以当前状态、绝对时间、剩余时域和采样成本为条件的共享继续值模型。我们证明精确价值函数和继续函数在成本上非减、凹且与时域相关的Lipschitz连续,并将其性质嵌入模型架构,推导出基于残差的价值与策略误差界。在受控高斯过程、时变非高斯过程及FordA发动机噪声时序基准测试中,与代表性逐点反向归纳求解器和调优静态停时规则对比,单个CC-AOS检查点在6个未见的FordA成本-时域组合上,均优于独立训练的凸函数学习方法,平均降低目标函数(终端风险+采样成本)15.75%,且平均性能持平于调优静态阈值。

原文摘要 · Abstract (English)

Finite-horizon optimal stopping is a central problem in early time-series classification, where a system must decide at each sequence prefix whether the expected benefit of another observation justifies its acquisition cost. Existing data-driven backward-induction methods typically solve each cost-horizon operating point separately, so changing operating conditions requires repeated optimization and separate model stacks, making continuous cost adaptation and multi-horizon deployment inefficient. We propose CC-AOS (Cost- and Horizon-Conditioned Amortized Optimal Stopping), a structured amortized solver for a family of finite-horizon stopping problems with continuous costs and multiple horizons. CC-AOS learns a shared continuation-value model conditioned on the current state, absolute time, remaining horizon, and acquisition cost through joint amortized fitted backward induction. We establish that the exact value and continuation functions are nondecreasing, concave, and horizon-dependently Lipschitz in cost, encode these properties in the model architecture, and derive residual-based bounds on value and policy errors. Experiments on controlled Gaussian and time-varying non-Gaussian processes and the FordA engine-noise time-series benchmark compare CC-AOS with representative per-operating-point backward-induction solvers and tuned static stopping rules. At six unseen FordA cost-horizon pairs, one CC-AOS checkpoint achieved a lower terminal-risk-plus-sampling-cost objective than independently fitted Convex Function Learning at all six pairs, with an average reduction of 15.75 percent, while matching the tuned static thresholds on average.

最优停时时序决策成本感知机器学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。