让机器人在不同光照下稳定工作,无需重新收集数据
RoHIL: Robust Human-in-the-Loop Robotic Reinforcement Learning Against Illumination Variations

- 用世界模型重生成光源变化下的视觉数据,保留真实动作和奖励
- 新方法使机器人在光照变化后性能显著提升,且原环境表现不变
- 适合需要跨场景部署的工业机器人,避免重复训练
人类在环强化学习系统在训练工作站上表现近乎完美,但当同一机器人移至数米外的工作站时,因光照变化导致视觉输入分布偏移而崩溃。重新收集演示并重新运行人机协同强化学习不适用于部署,而直接在光照变化数据上微调则引发灾难性遗忘。为此,我们提出RoHIL——一种无需额外真实机器人交互的离线微调框架。该框架包含:(i) 基于世界模型的图像重照明技术,可在多种虚拟HDRI环境下重生成源工作站轨迹的视觉流,保持动作与奖励真实;(ii) 光照保留回放(IRR)机制,通过交替插入重照明适应转换与原始光照保留转换,维持源工作站贝尔曼覆盖;(iii) 锚定贝尔曼-策略正则化,限制表示与策略从原始策略的漂移。在四个真实机器人操控任务中,罗希尔在显著跨工作站光照变化下显著提升迁移性能,而源工作站性能保持不变,无需为每个新环境重新采集数据或训练。
原文摘要 · Abstract (English)
Human-in-the-loop reinforcement learning systems achieve near-perfect success on the workstation where they are trained, but collapse when the same robot is moved to a workstation a few meters away due to shifts in the visual input distribution caused by new lamp positions and window light. Re-collecting demonstrations and re-running HIL on every workstation is incompatible with deployment, and naively fine-tuning on shifted-light data triggers catastrophic forgetting of the source workstation. To close this cross-domain gap, we present RoHIL, an offline fine-tuning framework that uses no extra real-robot interaction. RoHIL combines (i) a world-model-based image relighter that re-synthesises the visual stream of source-workstation trajectories under multiple virtual HDRI environments, leaving actions and rewards real; (ii) Illumination-Retention Replay (IRR), a data-level anti-forgetting mechanism that interleaves relit adaptation transitions with original-light retention transitions to preserve source-workstation Bellman coverage; and (iii) an anchored Bellman-actor regulariser that constrains representation and policy drift from the original source-workstation policy. Across four real-robot manipulation tasks under significant cross-workstation illumination variations, RoHIL substantially improves shifted-light performance where standard HIL-RL collapses, while preserving source-workstation performance, eliminating the need to re-collect data and retrain for every new workstation and environment. Project page: https://anonymous4365.github.io/RoHIL/
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。