提出在线部署后动态调整策略,应对实时通信中的未知干扰。
Streetwise Agents: Empowering Offline RL Policies to Outsmart Exogenous Stochastic Disturbances in RTC
- 部署后根据实时分布外区域特征动态修正策略
- 在带宽估计任务中性能提升约18%
- 适合需要高可靠性的实时通信系统
真实生产系统中在线探索与训练的困难限制了实时数据驱动决策的应用。最可行方案是基于有限轨迹样本采用离线强化学习。然而,部署后因外部因素临时或永久扰动决策过程的转移分布,导致策略失效与泛化误差,尤其在实时通信(RTC)等敏感领域尤为严重。本文解决因未见外部随机扰动引发的领域偏移下识别鲁棒动作的关键问题。由于无法在离线数据支持范围内学习对未见外部扰动具备鲁棒性的通用策略,我们提出一种新的部署后策略塑造方法(Streetwise),基于对分布外子空间的实时刻画进行条件化调整。该方法显著提升了实时通信中网络瓶颈带宽估计(BWE)的鲁棒性,并在标准离线RL基准测试中表现优异。大量实验结果表明,在部分场景下最终回报相较当前最优基线提升约18%。
原文摘要 · Abstract (English)
The difficulty of exploring and training online on real production systems limits the scope of real-time online data/feedback-driven decision making. The most feasible approach is to adopt offline reinforcement learning from limited trajectory samples. However, after deployment, such policies fail due to exogenous factors that temporarily or permanently disturb/alter the transition distribution of the assumed decision process structure induced by offline samples. This results in critical policy failures and generalization errors in sensitive domains like Real-Time Communication (RTC). We solve this crucial problem of identifying robust actions in presence of domain shifts due to unseen exogenous stochastic factors in the wild. As it is impossible to learn generalized offline policies within the support of offline data that are robust to these unseen exogenous disturbances, we propose a novel post-deployment shaping of policies (Streetwise), conditioned on real-time characterization of out-of-distribution sub-spaces. This leads to robust actions in bandwidth estimation (BWE) of network bottlenecks in RTC and in standard benchmarks. Our extensive experimental results on BWE and other standard offline RL benchmark environments demonstrate a significant improvement ($\approx$ 18% on some scenarios) in final returns wrt. end-user metrics over state-of-the-art baselines.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。