分钟级训练实现微机器人自主导航,突破传统强化学习耗时瓶颈。
Minute-Scale Training for Microrobot Navigation

- 构建千级血管环境的向量化仿真器,每秒生成19万次状态转移。
- 提出任务塑造正则化奖励框架,训练时间压缩至10分钟内。
- 支持零样本部署,适用于不同机器人类型和场景,适合快速原型开发。
微机器人在多种应用中具有巨大潜力,靶向导航是其基本需求。深度强化学习(DRL)近期成为实现全自主微机器人导航的强大范式,但现有方法对学习效率与效果关注不足,模型训练需数小时至数日,严重制约快速部署与参数优化。为此,我们提出一种学习框架,可在分钟内完成有效导航策略训练。该框架包含一个包含超10,000个虚拟血管环境的全向量化仿真器,通过并行处理动力学、类激光雷达感知及可行性检查,实现约190,000次状态转移/秒。为提升快速训练下的性能,提出任务塑造正则化(TSR)奖励机制:加速收敛,提升最终表现,降低动作波动至少33.7%,障碍物清除率提升至少2.1%。实验表明,该框架将训练时间缩短至10分钟以下,并支持跨不同微机器人类型与导航场景的零样本部署。整体上,该框架显著缩短设计周期,加速自主微机器人的实际应用。
原文摘要 · Abstract (English)
Microrobots hold significant potential for various applications, where targeted navigation is a basic requirement. Deep reinforcement learning (DRL) has recently emerged as a powerful paradigm for fully autonomous microrobot navigation. Yet, current DRL-based approaches pay limited attention to learning efficiency and effectiveness, requiring hours to days for model training. Consequently, this impedes both rapid practical deployment and parameter optimization. To address these challenges, we present a learning framework that enables effective microrobot navigation policies to be trained within minutes. In the proposed framework, we develop a fully vectorized simulator with more than 10,000 artificial vascular environments, parallelizing dynamics, LiDAR-inspired perception, and feasibility checks across thousands of environments to achieve roughly 190,000 transitions per second. To achieve effectiveness in fast training, we propose a task-shaping-regularization (TSR) reward framework. The TSR framework accelerates convergence, improves final performance, reduces action variation by at least 33.7%, and increases obstacle clearance by at least 2.1% across all evaluated scenarios. Results show that the proposed learning framework reduces training time to under 10 minutes, while supporting zero-shot deployment across distinct microrobot types and navigation scenarios. Collectively, this framework can substantially shorten the design loop and accelerate the deployment of autonomous microrobots.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。