用马氏距离检测异常束流分布,实现强化学习与模型无关控制的自动切换。
Mahalanobis-Guided Latent OOD Detection for Hybrid ES-DRL Control in Time-Varying Systems

- 在变分自编码器隐空间用马氏距离检测测试时的分布外数据
- 在粒子加速器控制中成功识别未训练过的束流分布并触发控制器切换
- 方法可解释,适合高安全要求的时变系统控制场景
本文研究了在非线性时变系统中,基于马氏距离引导的潜在空间分布外(OOD)检测,用于强化学习(RL)控制器在测试时的切换。尽管强化学习控制器能在训练分布内快速控制高维系统,但在时变动态导致未见观测时性能会下降。本文采用混合型ES-DRL控制器:强化学习提供快速的分布内动作,有界极值搜索(ES)则在分布外操作下提供无需模型的鲁棒控制。核心挑战在于何时切换。我们对分布内束流轮廓观测训练一个变分自编码器(VAE),并在测试时利用其隐空间中的马氏距离检测分布外束流轮廓。该分布外判定作为二元开关,选择使用强化学习或极值搜索控制器。我们在安全关键的粒子加速器控制中评估该方法,在此场景中,空间磁铁运动产生了训练时未见的分布外束流轮廓。通过可视化VAE隐空间,发现所提方法能有效识别该分布外情形,并为混合控制器的切换提供可解释信号。
原文摘要 · Abstract (English)
In this paper, we study Mahalanobis-guided latent out-of-distribution (OOD) detection for test-time RL controller switching in nonlinear time-varying systems. RL controllers can quickly control high-dimensional systems within the training distribution, but their performance can degrade when time-varying dynamics produce unseen observations. We consider a combined ES--DRL controller, where RL provides fast in-distribution actions and bounded extremum seeking (ES) provides robust model-independent control under OOD operation. The key challenge is deciding when to switch. We train a variational autoencoder (VAE) on in-distribution beam-profile observations and use Mahalanobis distance in the VAE latent space to detect OOD beam profiles at test time. This OOD decision sets a binary switch that selects either the RL controller or the ES controller. We evaluate the approach in safety-critical particle accelerator control. In this setting, spatial magnet motion creates OOD beam profiles that were not seen during RL training. Visualization of the VAE latent space shows that the proposed method identifies this OOD scenario and provides an interpretable signal for switching between RL and ES in the combined controller.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。