用强化学习直接优化场景树,让预测更贴合实际控制效果。
Control-Oriented Scenario Tree Construction through Reinforcement Learning

- 通过注意力机制+强化学习,把预测样本分配到树叶节点
- 在电池套利任务中,收益比传统方法最高提升12.3%
- 适合需要应对不确定性的工业控制系统设计者
多阶段随机模型预测控制(MPC)通过构建场景树来处理不确定性,该树是未来结果的有限分支近似,由采样预测生成。传统方法聚焦于匹配底层概率分布(如基于Wasserstein的距离缩减),但分布精度提升并不一定带来更好的控制性能。本文提出一种以控制为导向的方法,直接从场景树对下游决策的影响中学习其构建方式。固定树结构后,将场景分配至叶子节点建模为序列分配问题,采用基于注意力的策略参数化,并以闭环控制收益为目标进行强化学习训练。训练通过非对称批评器稳定,利用实际发生的未来轨迹。在风险厌恶型电池套利问题上评估表明,无论采样集大小如何,所学构造均持续获得最高收益,显著优于经典前向/后向缩减方法及确定性等价(单轨迹预测)控制。所学策略在困难实例中表现出更强鲁棒性,尾部风险特征更优。分析显示,本方法构建出紧凑、选择性分叉的结构,捕捉高影响事件的同时,多数路径保持近乎确定。这说明场景树的价值取决于其所支持的决策,提供了一种仅依赖闭环控制信号训练场景树构造器的有效框架。
原文摘要 · Abstract (English)
Multistage stochastic model predictive control (MPC) handles uncertainty by optimizing over a scenario tree, a finite branching approximation of future outcomes constructed from sampled forecasts. To build such a tree, conventional methods focus on matching the underlying probability distribution---e.g., via Wasserstein-based scenario reduction---but improved distributional accuracy does not necessarily yield better control performance. We propose a control-oriented approach that learns scenario tree construction directly from its impact on downstream decisions. Fixing the tree topology, we formulate tree construction as a sequential assignment of sampled scenarios to leaves. This assignment is parameterized by an attention-based policy over the scenario set and trained using reinforcement learning, with closed-loop control profit as the objective. Training is stabilized by an asymmetric critic that leverages realized future trajectories. We evaluate the method on a risk-averse battery arbitrage problem. Across a range of forecast set sizes, the learned construction consistently achieves the highest profit, outperforming classical forward and backward reduction methods and certainty-equivalent (single-trajectory forecast) control. The learned policy also exhibits greater robustness on challenging instances, consistently demonstrating better tail-risk characteristics. Analysis of the resulting trees indicates that our method constructs compact, selectively branching structures that capture high-impact events while keeping most trajectories nearly deterministic. These findings highlight that the value of a scenario tree depends critically on the decisions it supports, and provide an effective framework to train scenario tree constructors merely based on the closed-loop control optimization signal.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。