提醒强化学习研究者区分‘解仿真环境’和‘用仿真做部署代理’两种目标。
Position: RL Researchers Need to Distinguish Between Solving Simulators and Using Simulators as a Proxy
- 明确区分两类仿真使用场景:解仿真器与用仿真器模拟真实部署。
- 指出混淆二者会导致评估指标错位和算法选择偏差。
- 适合关注实验设计严谨性的RL研究者阅读。
强化学习(RL)研究的一个目标是理解通用的序贯决策问题,常以基准仿真环境作为部署场景的学习代理。然而在实验中,追求仿真环境中的高分可能演变为只专注于‘解仿真器’,导致研究者采用仅针对仿真优化的策略,而非为实际部署而学习。虽然‘解仿真器’本身值得研究,但这是与真实部署完全不同的问题。本文主张研究人员必须区分两种仿真使用方式:一是纯粹解决仿真环境,二是将仿真作为部署学习的代理。我们从代理对仿真使用的约束、适用算法及评估指标等角度说明二者本质差异,并通过实例和简单实验揭示混淆两者的潜在问题和误导性结论。本文呼吁社区在工作中清晰界定仿真用途,推动对不同场景下最佳实证实践的深入讨论。
原文摘要 · Abstract (English)
One goal in reinforcement learning (RL) research is to understand general-purpose sequential decision-making, using benchmark simulators as a proxy for learning in deployment settings. When running experiments, however, the goal of achieving high performance in the simulator can mutate into focusing exclusively on solving the simulator. To achieve high scores, researchers may adopt solutions exclusively meant for solving simulators, rather than learning while the agent is deployed outside a simulator. Solving simulators is also worthy of investigation, but it is a fundamentally different RL research question. In this paper, we argue that RL researchers need to distinguish between two use cases of simulators: solving simulators and using simulators as a proxy for learning in deployment. We first discuss how these two use-cases are importantly different, in terms of constraints on how the agent can use the simulator, which algorithms are appropriate, and which evaluation metrics are appropriate. We then highlight several issues and misleading conclusions that can occur by not making the distinction between these two settings clear, supported with examples and simple experiments. This work is a call to the community to begin clearly distinguishing how they are using simulators in their work, hopefully sparking further discussion on which empirical practices work best in each setting.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。