arXiv:2606.19357cs.ROcs.AI2026-06被引 1

用低成本硬件搭建可长期运行的物理强化学习实验平台

Physical Atari: A Robust and Accessible Platform for Real-time Reinforcement Learning on Robots

论文配图:Physical Atari: A Robust and Accessible Platform for Real-time Reinforcement Learning on Robots
图 1 · 摘自论文原文
  • 用轴承驱动机械臂精准操控游戏手柄,减少磨损
  • 系统持续运行数周无故障,成本低于1000美元
  • 适合研究机器人强化学习及分布偏移影响

我们构建了名为Robotroller的机器人,可操作Atari CX40+控制器,并开发了Atari Devbox设备,将Arcade Learning Environment的游戏画面和奖励信号渲染到屏幕上。Robotroller与Atari Devbox配合现成摄像头和台式机,构成一个可在真实世界中研究强化学习算法的系统,称为Physical Atari。本文详细说明了使该系统具备鲁棒性与可访问性的关键设计:为减少磨损,机器人所有运动均通过轴承实现;同时编写高频监控软件,实时检测舵机状态并干预以降低应力。为提升可访问性,采用廉价现成组件,且大部分零件可用消费级3D打印机制造。Physical Atari总成本低于1000美元,已成功支持数周不间断强化学习实验。我们验证了强化学习算法可直接在机器人上训练,并发现学习与部署间的小分布偏移会显著降低策略性能,强调了设备端自适应对机器人强性能的重要性。

原文摘要 · Abstract (English)

We built a robot called the Robotroller that actuates an Atari CX40+ controller and a device called the Atari Devbox that renders the game frame and the reward signal from the Arcade Learning Environment on a screen. The Robotroller and the Atari Devbox, together with an off-the-shelf camera and a desktop computer, constitute a system that can be used to study reinforcement learning algorithms in the physical world. We call the full system Physical Atari. In this paper, we detail the key decisions that make Physical Atari a robust and accessible platform. To make the system robust, we designed the Robotroller so that all movement is done through bearings, which reduces wear. Additionally, we wrote software that monitors the state of the servos at a high frequency and intervenes to limit stress. To make the system accessible, we used affordable off-the-shelf components and parts that can be manufactured using consumer 3D printers. Physical Atari can be built for under $1,000 and has been used for weeks of non-stop reinforcement learning experiments without any mechanical failures. We used it to validate that reinforcement learning algorithms can learn directly on robots and show that even small distribution shifts between learning and deployment can significantly degrade the performance of policies. Our results underscore the importance of on-device adaptation for strong performance on robots.

强化学习机器人可复现低成本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。