用单阶段方法让机器人在未知环境下稳定运行
Out-of-Distribution Generalization with a SPARC: Racing 100 Unseen Vehicles with a Single Policy
- 设计单阶段适配框架,无需分步训练
- 在多个未见场景中表现稳定,泛化能力强
- 适合无人车、机器人等需跨环境控制的场景
机器人与控制领域面临的重要挑战是适应未见过的环境。本文聚焦于上下文强化学习,即智能体在不同上下文环境中行动,如自动驾驶汽车或四足机器人需在训练时未遇的地形或天气条件下运行。我们解决测试时无法获取显式上下文信息条件下的分布外(OOD)泛化问题。现有方法采用分阶段训练上下文编码器和历史适应模块,虽有效但实现与训练复杂。本文简化流程,提出SPARC:单阶段鲁棒控制适配框架。在高保真赛车模拟器Gran Turismo 7及受风扰动的MuJoCo环境中验证,SPARC展现出可靠的分布外泛化能力。
原文摘要 · Abstract (English)
Generalization to unseen environments is a significant challenge in the field of robotics and control. In this work, we focus on contextual reinforcement learning, where agents act within environments with varying contexts, such as self-driving cars or quadrupedal robots that need to operate in different terrains or weather conditions than they were trained for. We tackle the critical task of generalizing to out-of-distribution (OOD) settings, without access to explicit context information at test time. Recent work has addressed this problem by training a context encoder and a history adaptation module in separate stages. While promising, this two-phase approach is cumbersome to implement and train. We simplify the methodology and introduce SPARC: single-phase adaptation for robust control. We test SPARC on varying contexts within the high-fidelity racing simulator Gran Turismo 7 and wind-perturbed MuJoCo environments, and find that it achieves reliable and robust OOD generalization.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。