低功耗神经脉冲扩散策略提升航天器内机械臂操作成功率
L-SDPPO: Policy Optimization of Spiking Diffusion Policy for Intra-vehicular Robotic Manipulation

- 用脉冲神经网络优化扩散策略,降低计算能耗
- 在5项舱内任务中成功率达92%以上,能耗降低40%
- 适合微重力环境下的高可靠性机器人控制
航天器内的舱内机器人可减轻宇航员负担、提升任务效率。现有研究多采用深度学习实现复杂环境下的精确控制,但物体在无重力环境下漂移无约束,需应对复杂的多模态动作分布。扩散策略(DP)虽能建模复杂动作,但其迭代采样过程能耗过高,不适用于航天器有限的供电条件。为此,我们提出低能耗舱内机器人操作框架L-SDPPO,通过强化学习优化脉冲扩散策略(SDP)。同时,为解决微重力下动态时空特征感知不足的问题,提出状态依赖延迟注入(SDLI)机制,模拟生物神经延迟以动态调节输入信息时机。在五项典型舱内日常任务(如舱门开启、精密容器封盖)上的评估表明,本方法在成功率和能耗方面均优于当前最优机器人操控方法,验证了其作为可行舱内机器人操作方案的潜力。
原文摘要 · Abstract (English)
Intra-vehicular robots in spacecraft help reduce astronaut workload and improve mission efficiency. Recent research focuses on using deep learning methods to achieve the acute control required for operations in these complex environments. However, objects exhibit unpredictable, unconstrained drift without gravitational damping. These factors demand robustness against complex multimodal action distributions. Diffusion policies (DP) can model these complex actions, but their iterative sampling process consumes too much energy for the limited power budgets of spacecraft. We therefore propose a low-energy intra-vehicular robotic manipulation framework, L-SDPPO, in which the Spiking Diffusion Policy (SDP) is optimized with a reinforcement learning (RL) algorithm. Furthermore, to address the insufficient perception of dynamic spatiotemporal features in microgravity, we propose the statedependent latency injection (SDLI) mechanism, which mimics biological neural delays to dynamically regulate the timing of input information. Evaluation on five representative intra-vehicular daily tasks (e.g., hatch opening and precision container capping) shows that our method consistently achieves higher success rates and lower energy consumption, compared to the state-of-the-art robotic manipulation methods. These results demonstrate our method is a viable intra-vehicular robotic manipulation method.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。