让生成流网络在机器人任务中无需预训练就能自适应,提升稳定性和故障应对能力。
WINFlowNets: Warm-up Integrated Networks Training of Generative Flow Networks for Robotics and Machine Fault Adaptation
- 通过热身阶段启动检索网络,再联合训练流网络与检索网络。
- 在模拟机器人环境中平均奖励更高,训练更稳定,且故障适应性更强。
- 适合动态环境或样本稀缺场景,如故障机器人快速恢复控制。
连续场景生成流网络(CFlowNets)通过流网络和检索网络学习随机策略,在序列决策任务中表现出色,相比先进强化学习算法效率更高。然而其在机器人控制中的实际应用受限于对检索网络预训练的依赖,而在动态环境中,预训练数据可能难以获取或不具代表性。本文提出WINFlowNets,一种新型CFlowNets框架,实现流网络与检索网络的联合训练。该方法首先通过热身阶段初始化检索网络以启动策略,随后采用共享训练架构和共享回放缓冲区进行联合优化。在模拟机器人环境中的实验表明,WINFlowNets在平均奖励和训练稳定性上均优于CFlowNets及现有先进RL算法。此外,该方法在故障环境下表现出强适应能力,适用于仅有限样本数据时需快速适应的任务。这些结果表明,WINFlowNets具备在动态且易出故障的机器人系统中部署的潜力,尤其适用于传统预训练或低样本数据收集不可行的场景。
原文摘要 · Abstract (English)
Generative Flow Networks for continuous scenarios (CFlowNets) have shown promise in solving sequential decision-making tasks by learning stochastic policies using a flow and a retrieval network. Despite their demonstrated efficiency compared to state-of-the-art Reinforcement Learning (RL) algorithms, their practical application in robotic control tasks is constrained by the reliance on pre-training the retrieval network. This dependency poses challenges in dynamic robotic environments, where pre-training data may not be readily available or representative of the current environment. This paper introduces WINFlowNets, a novel CFlowNets framework that enables the co-training of flow and retrieval networks. WINFlowNets begins with a warm-up phase for the retrieval network to bootstrap its policy, followed by a shared training architecture and a shared replay buffer for co-training both networks. Experiments in simulated robotic environments demonstrate that WINFlowNets surpasses CFlowNets and state-of-the-art RL algorithms in terms of average reward and training stability. Furthermore, WINFlowNets exhibits strong adaptive capability in fault environments, making it suitable for tasks that demand quick adaptation with limited sample data. These findings highlight WINFlowNets' potential for deployment in dynamic and malfunction-prone robotic systems, where traditional pre-training or sample inefficient data collection may be impractical.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。