解决强化学习在真实世界落地中的样本效率与系统持续优化难题。
Reinforcement Learning in the Real World: A Survey of Statistical Challenges and Future Directions
- 将实际应用拆解为在线学习、离线分析和迭代部署三阶段
- 提出提升部署中样本效率与离线数据分析价值的方法
- 面向医疗、机器人等动态场景,推动可持续优化的RL系统
强化学习(RL)在游戏、机器人、在线广告、公共卫生及自然语言处理等领域实现了显著进展。然而,理论研究与实际部署之间仍存在显著差距。两大核心挑战是:环境交互受限导致样本稀缺;环境本身持续变化需频繁重设计与重部署。本文将实际应用归纳为三个环节:部署中的在线学习与优化、部署后或间歇期的离线分析、以及反复迭代的部署-再部署循环以实现系统持续改进。综述了近年来在提升在线部署样本效率、增强离线分析数据利用率、设计可持续改进的部署序列等方面的统计方法进展,并展望了面向实际应用的未来研究方向。
原文摘要 · Abstract (English)
Reinforcement learning (RL) has achieved remarkable success in real-world decision-making across diverse domains, including gaming, robotics, online advertising, public health, and natural language processing. Despite these advances, a substantial gap remains between RL research and its deployment in many practical settings. Two recurring challenges often underlie this gap. First, many settings offer limited opportunity for the agent to interact extensively with the target environment due to practical constraints. Second, many target environments often undergo substantial changes, requiring redesign and redeployment of RL systems (e.g., advancements in science and technology that change the landscape of healthcare delivery). Addressing these challenges and bridging the gap between basic research and application requires theory and methodology that directly inform the design, implementation, and continual improvement of RL systems in real-world settings. In this paper, we frame the application of RL in practice as a three-component process: (i) online learning and optimization during deployment, (ii) post- or between-deployment offline analyses, and (iii) repeated cycles of deployment and redeployment to continually improve the RL system. We provide a narrative review of recent advances that address the statistical challenges arising across these three components, including methods for enhancing sample efficiency during online deployment, maximizing data utility for post- or between-deployment inference, and designing sequences of deployments for continual improvement. We also outline future research directions in RL that are use-inspired -- aiming for impactful application of RL in practice.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。