arXiv:2510.05147cs.SEcs.LG2025-10

用强化学习动态分配测试资源,自动适应故障率变化。

Adaptive Reinforcement Learning for Dynamic Configuration Allocation in Pre-Production Testing

  • 将配置分配建模为序列决策问题,结合模拟与实时反馈优化
  • 在仿真中接近最优表现,显著优于静态基线方法
  • 适合需要快速响应环境变化的软件测试与资源调度场景

现代软件系统可靠性依赖于在高度异构且不断演化的环境中进行严格的预生产测试。由于全面评估不可行,实践者需在有限测试资源下分配到不同配置中,而这些配置的故障概率可能随时间漂移。现有组合优化方法静态、随意,难以应对非平稳环境。本文提出一种新型强化学习框架,将配置分配重构为序列决策问题。该方法首次结合Q-learning与混合奖励设计,融合模拟结果与实时反馈,兼顾样本效率与鲁棒性。此外,我们开发了自适应在线-离线训练机制,使智能体能快速追踪突发的概率变化,同时保持长期稳定性。大量仿真表明,本方法持续优于静态和基于优化的基线,逼近理想性能(oracle performance)。这项工作确立了强化学习在自适应配置分配中的新范式,超越传统方法,在动态测试与资源调度领域具有广泛适用性。

原文摘要 · Abstract (English)

Ensuring reliability in modern software systems requires rigorous pre-production testing across highly heterogeneous and evolving environments. Because exhaustive evaluation is infeasible, practitioners must decide how to allocate limited testing resources across configurations where failure probabilities may drift over time. Existing combinatorial optimization approaches are static, ad hoc, and poorly suited to such non-stationary settings. We introduce a novel reinforcement learning (RL) framework that recasts configuration allocation as a sequential decision-making problem. Our method is the first to integrate Q-learning with a hybrid reward design that fuses simulated outcomes and real-time feedback, enabling both sample efficiency and robustness. In addition, we develop an adaptive online-offline training scheme that allows the agent to quickly track abrupt probability shifts while maintaining long-run stability. Extensive simulation studies demonstrate that our approach consistently outperforms static and optimization-based baselines, approaching oracle performance. This work establishes RL as a powerful new paradigm for adaptive configuration allocation, advancing beyond traditional methods and offering broad applicability to dynamic testing and resource scheduling domains.

强化学习测试优化动态调度

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。