用软专家混合提升强化学习在控制器合成中的探索鲁棒性
Robust Exploration in Directed Controller Synthesis via Reinforcement Learning with Soft Mixture-of-Experts
- 通过先验置信度门控融合多个强化学习专家
- 在航空交通基准上扩展了可解参数空间,性能更稳定
- 适合需要高鲁棒性的实时控制系统设计者
在线定向控制器合成(OTF-DCS)通过增量式探索系统来缓解状态空间爆炸问题,其关键依赖于高效的探索策略。近期基于强化学习(RL)的方法能够从少量训练实例中学习探索策略,并实现出色的零样本泛化能力。然而,其根本局限在于方向性泛化:由于训练随机性和轨迹依赖偏差,RL策略仅在特定领域参数空间区域内表现良好,其他区域则易失效。为此,本文提出软专家混合(Soft-MoE)框架,通过先验置信度门控机制融合多个RL专家,并将这种方向性缺陷视为互补的专业化能力。在航空交通基准上的评估表明,Soft-MoE显著扩展了可解参数空间,相比单个专家提升了整体鲁棒性。
原文摘要 · Abstract (English)
On-the-fly Directed Controller Synthesis (OTF-DCS) mitigates state-space explosion by incrementally exploring the system and relies critically on an exploration policy to guide search efficiently. Recent reinforcement learning (RL) approaches learn such policies and achieve promising zero-shot generalization from small training instances to larger unseen ones. However, a fundamental limitation is anisotropic generalization, where an RL policy exhibits strong performance only in a specific region of the domain-parameter space while remaining fragile elsewhere due to training stochasticity and trajectory-dependent bias. To address this, we propose a Soft Mixture-of-Experts framework that combines multiple RL experts via a prior-confidence gating mechanism and treats these anisotropic behaviors as complementary specializations. The evaluation on the Air Traffic benchmark shows that Soft-MoE substantially expands the solvable parameter space and improves robustness compared to any single expert.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。