实验证明强化学习编排在真实环境中的表现被高估,需更严格的评估标准。
Incentives and Evidence in Learned Service Orchestration
- 用预注册测试验证三种主流RL编排系统在扰动下的表现
- 多数预测的性能下降未出现,部分结果无法复现
- 建议引入生产级对比器和可重复的运行指标
过去十余年,强化学习在服务编排领域的研究持续不断,但尚未大规模投入生产。通常解释为学习控制器在延迟和噪声遥测、负载变化及不可控租户环境下性能下降。我们检验现有证据是否支持这一说法。对三种具有代表性的基于RL的编排系统(资源分配、DAG调度、自动扩缩)进行评估,采用预注册预测,考察其在生产相关扰动下的性能退化,并使用族系误差校正进行配对推理。结果显示,大多数预测的性能反转并未发生。诊断分析表明,这些结果往往源于比较器崩溃、数据集局限或评估设计问题,而非真实耐受性证据。一种看似在观测延迟下有40倍优势的结果,实际上远低于公开报告值;另一广泛引用的结果无法从发布代码中复现,最可复现的差距也远小于原结果。结论随扰动幅度和评估模式变化而反转。基于这些发现及文献中的普遍模式,我们识别出制度性问题:发表与评审激励偏好对便捷对照组的基准提升,即使这些提升无法反映实际部署表现。问题不单是技术性的,更是制度性的,需要生产级对照组、注册扰动模型、独立运营指标以及鼓励可复现操作证据的发表标准。否则,文献虽增长,却无法确认学习是否真正改进了编排。
原文摘要 · Abstract (English)
Reinforcement learning for service orchestration has been the subject of sustained research for over a decade, yet it is not used in production at scale. The usual explanation is that learned controllers degrade under delayed and noisy telemetry, workload shifts, and uncontrolled tenants. We test whether existing evidence supports that explanation. We evaluate three highly influential RL-based orchestration systems spanning resource allocation, DAG scheduling, and autoscaling, using pre-registered predictions about comparative degradation under production-relevant perturbations and paired inference with family-wise error correction. Across the tests, most predicted performance reversals do not occur. Diagnostic analyses show that these outcomes often reflect comparator collapse, artefact limitations, or evaluation choices rather than evidence that learned controllers tolerate the perturbations. One apparent advantage under observation lag is roughly fortyfold compared to a Kubernetes HPA-equivalent controller. Another widely cited result cannot be reconstructed from its released artefact, and the strongest reproducible margin is far smaller than the published results. Conclusions also reverse under changes in perturbation magnitude and evaluation mode. Based on these results and broader patterns in the literature, we identify an institutional problem. Publication and review incentives favour benchmark gains against convenient comparators, even when those gains provide little evidence of deployment performance. We argue that the problem is not solely technical. Rather, it is institutional, so learned orchestration needs production-grade comparators, registered perturbation models, separate operational metrics, and publication criteria that reward reproducible operational evidence. Without these changes, the literature can grow without establishing whether learning improves orchestration.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。