让控制策略持续保持鲁棒性,应对系统参数随时间变化的不确定性。
Robust Control under Stationary Ambiguity

- 设计状态相关的动态不确定性模拟器,使参数不确定性不随时间衰减。
- 在金融对冲任务中,策略在真实市场数据上表现更稳定,鲁棒性更强。
- 适合需要长期应对未知环境变化的控制场景,如金融市场、自适应系统。
在仿真中优化的控制策略在现实系统中表现不佳,常因仿真参数 $x$ 由有限数据估计而引入不确定性,但该不确定性未在仿真中体现。传统方法通过随机抽样 $x$ 模拟轨迹,使策略初始具备泛化能力,但随着系统观测积累,策略逐渐推断出 $x$ 并失去鲁棒性。为解决此问题,本文提出「平稳不确定性」(stationary ambiguity):让不确定性随系统状态变化但不随时间系统性衰减。我们构建满足该性质的模拟器,在对冲任务中验证了策略能持续保持对潜在因素的鲁棒性,真实市场数据表现优异。该原则可指导仿真建模、参数随机化及初始化策略,适用于受外生随机过程驱动、潜在结构随时间变化的序列控制问题。
原文摘要 · Abstract (English)
Control policies optimized in simulation can perform poorly in the real system when the parameters $x$ of the simulator are estimated from limited data but the resulting parameter uncertainty is not represented inside the simulation. A common way to incorporate such ambiguity is to simulate each trajectory of the system under a randomly drawn value for $x$. Since the policy cannot observe the drawn value, it must initially choose controls that perform well across many possible parameter values. However, if the policy progressively observes the system, it can often gradually infer the value of $x$, so that ambiguity vanishes. Over time, the policy then specializes to its estimate of $x$ and loses its robustness. This is undesirable in many real systems, where latent factors are expected to shift. In financial markets, for example, a policy hedging a derivative payoff should remain robust to changes in the volatility regime. To induce such continual robustness, we propose training policies in simulators where ambiguity varies with the system's state but does not systematically decay over time. We formalize this requirement as stationary ambiguity: the simulator should induce a stationary filter process over the latent state. We show how to construct such simulators and demonstrate, on hedging problems, that policies trained under stationary ambiguity preserve robustness to latent factors over time, leading to strong performance on real market data. As a modeling principle, stationary ambiguity informs many simulator design decisions: which models make realistic simulators, how their parameters should be randomized, and how simulator and policy should be initialized. While our experiments focus on hedging, stationary ambiguity may also be useful for other sequential control problems driven by exogenous stochastic processes with shifting latent structure.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。