arXiv:2603.18396cs.LGcs.RO2026-03被引 1

分离随机与认知不确定性,提升公交调度的稳定性和鲁棒性。

RE-SAC: Disentangling aleatoric and epistemic risks in bus fleet control: A stable and robust ensemble DRL approach

  • 用积分概率度量正则化区分两类不确定性
  • 在真实双向公交走廊中累计奖励提升约27%
  • 适合高波动交通环境下需要可靠决策的研究者

公交调度因交通和客流的随机性而困难。尽管深度强化学习有潜力,但标准演员-评论家算法在波动环境中易出现价值函数不稳定。其关键原因在于混淆了两类不确定性:不可消除的随机噪声(aleatoric)和数据不足导致的认知不确定性(epistemic)。将二者合并会引发噪声状态下的价值低估,导致策略崩溃。我们提出稳健集成软演员-评论家(RE-SAC)框架,显式解耦这两类不确定性。通过基于积分概率度量(IPM)的权重正则化,对评论家网络进行抗随机风险建模,提供平滑解析下界,无需昂贵的内层扰动。针对认知风险,采用多样化Q集合惩罚稀疏区域的过度自信估值。该双机制防止集成方差误将噪声识别为数据缺口,此失败模式已在消融实验中验证。在真实双向公交走廊仿真中,RE-SAC实现最高累计奖励(约 -0.4e6),优于原始SAC(-0.55e6)。马氏距离罕见性分析表明,RE-SAC在罕见分布外状态中,奥拉克Q值估计误差降低最多达62%(平均绝对误差1647 vs. 4343),展现高交通变异性下的卓越鲁棒性。

原文摘要 · Abstract (English)

Bus holding control is challenging due to stochastic traffic and passenger demand. While deep reinforcement learning (DRL) shows promise, standard actor-critic algorithms suffer from Q-value instability in volatile environments. A key source of this instability is the conflation of two distinct uncertainties: aleatoric uncertainty (irreducible noise) and epistemic uncertainty (data insufficiency). Treating these as a single risk leads to value underestimation in noisy states, causing catastrophic policy collapse. We propose a robust ensemble soft actor-critic (RE-SAC) framework to explicitly disentangle these uncertainties. RE-SAC applies Integral Probability Metric (IPM)-based weight regularization to the critic network to hedge against aleatoric risk, providing a smooth analytical lower bound for the robust Bellman operator without expensive inner-loop perturbations. To address epistemic risk, a diversified Q-ensemble penalizes overconfident value estimates in sparsely covered regions. This dual mechanism prevents the ensemble variance from misidentifying noise as a data gap, a failure mode identified in our ablation study. Experiments in a realistic bidirectional bus corridor simulation demonstrate that RE-SAC achieves the highest cumulative reward (approx. -0.4e6) compared to vanilla SAC (-0.55e6). Mahalanobis rareness analysis confirms that RE-SAC reduces Oracle Q-value estimation error by up to 62% in rare out-of-distribution states (MAE of 1647 vs. 4343), demonstrating superior robustness under high traffic variability.

强化学习公交调度不确定性建模鲁棒控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。