arXiv:2602.19244cs.AIcs.LG2026-02

用软专家混合提升强化学习在控制器合成中的探索鲁棒性

Robust Exploration in Directed Controller Synthesis via Reinforcement Learning with Soft Mixture-of-Experts

  • 通过先验置信度门控融合多个强化学习专家
  • 在航空交通基准上扩展了可解参数空间,性能更稳定
  • 适合需要高鲁棒性的实时控制系统设计者

在线定向控制器合成(OTF-DCS)通过增量式探索系统来缓解状态空间爆炸问题,其关键依赖于高效的探索策略。近期基于强化学习(RL)的方法能够从少量训练实例中学习探索策略,并实现出色的零样本泛化能力。然而,其根本局限在于方向性泛化:由于训练随机性和轨迹依赖偏差,RL策略仅在特定领域参数空间区域内表现良好,其他区域则易失效。为此,本文提出软专家混合(Soft-MoE)框架,通过先验置信度门控机制融合多个RL专家,并将这种方向性缺陷视为互补的专业化能力。在航空交通基准上的评估表明,Soft-MoE显著扩展了可解参数空间,相比单个专家提升了整体鲁棒性。

原文摘要 · Abstract (English)

On-the-fly Directed Controller Synthesis (OTF-DCS) mitigates state-space explosion by incrementally exploring the system and relies critically on an exploration policy to guide search efficiently. Recent reinforcement learning (RL) approaches learn such policies and achieve promising zero-shot generalization from small training instances to larger unseen ones. However, a fundamental limitation is anisotropic generalization, where an RL policy exhibits strong performance only in a specific region of the domain-parameter space while remaining fragile elsewhere due to training stochasticity and trajectory-dependent bias. To address this, we propose a Soft Mixture-of-Experts framework that combines multiple RL experts via a prior-confidence gating mechanism and treats these anisotropic behaviors as complementary specializations. The evaluation on the Air Traffic benchmark shows that Soft-MoE substantially expands the solvable parameter space and improves robustness compared to any single expert.

强化学习控制器合成专家混合

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。