用不到400美元搭建真实AIoT环境,测试强化学习的仿真到现实差距。
Measure the Sim-to-Real Gap: Designing an Affordable Real-World Benchmark Platform for Reinforcement Learning in AIoT Systems

- 用廉价硬件模拟键盘输入,让边缘设备在真实环境中玩视频游戏
- 仿真训练模型在真实环境性能下降1160%,暴露巨大仿真-现实差距
- 真实世界训练达人类水平49%,证明实际部署强化学习可行
强化学习常用于提升自主系统性能,包括智能物联网(AIoT)。然而,在真实环境中进行强化学习的试错过程成本高且存在风险。因此,多数研究仍依赖仿真环境,这带来了仿真到现实迁移的挑战。评估算法鲁棒性和仿真-现实差距是提升真实世界强化学习性能的关键前提。目前,机器人领域已有同步仿真与物理平台,但针对AIoT的通用仿真-现实基准平台尚不存在。为此,我们构建了一个低成本的真实世界AIoT平台,用于研究强化学习在AIoT中的应用。该平台使用成本低于400美元的商用组件,包含两台计算机:一个部署在边缘设备上的智能体通过硬件模拟键盘,接收视觉输入并在远程主机上玩视频游戏。由于目标是最大化游戏得分,该系统天然规避了真实部署的安全风险。实验结果显示,仿真训练的智能体在真实部署后性能较人类水平下降1160%,表明显著的仿真-现实差距。直接在真实环境中使用深度Q网络(DQN)训练,经过1000万步后达到约人类水平49%的性能,证明在真实条件下实施强化学习的可行性。这些结果表明,所提出的仿真-现实基准平台为真实世界AIoT系统中的强化学习提供了坚实的基础,可用于定性与定量评估。
原文摘要 · Abstract (English)
Reinforcement learning (RL) is commonly employed to enhance the performance of autonomous systems, including the Autonomous Internet of Things (AIoT). However, the trial-and-error nature of RL, when conducted in real-world environments, is costly and hazardous in some scenarios. Consequently, the majority of RL research is conducted in simulation. This reliance introduces challenges related to the Sim-to-Real transferability. Evaluating the Sim-to-Real algorithmic robustness and the Sim-to-Real gap is a critical prerequisite for research aimed at improving RL performance in the real world. Therefore, industries such as robotics have developed concurrent simulation and physical platforms to facilitate this research. However, a universal Sim-to-Real benchmark platform for AIoT does not currently exist. To address these concerns, we developed a real-world AIoT platform for studying RL in AIoT. On this platform, an agent deployed on an edge device plays video games on a separate host computer via a hardware-emulated keyboard, guided by vision input. This platform uses commercially available components costing less than USD 400, together with two computers. Because the system's objective is game score maximization, it inherently mitigates safety risks associated with real-world RL deployments. Experimental results show the simulation-trained agent suffers a 1160% performance degradation relative to the human-level performance after real-world deployment, indicating a significant Sim-to-Real gap. Direct real-world training using the deep Q-network (DQN) algorithm achieves approximately 49% of human-level performance after 10 million training steps, demonstrating the feasibility of RL under real-world conditions. These results suggest that the proposed Sim-to-Real benchmark platform provides a substantial foundation for qualitative and quantitative evaluations of RL in real-world AIoT systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。