用强化学习动态触发优化,让软机械臂控制更省算力
Deep Reinforcement Learning-Enhanced Event-Triggered Data-Driven Predictive Control for a 3D Cable-Driven Soft Robotic Arm

- 用无模型强化学习判断何时调用控制器,减少冗余计算
- 仿真中优化频率降低66%,硬件实验减少34%且精度相当
- 适合资源受限的实时软体机器人系统,可零样本迁移
软体机器人因非线性与时变动力学难以控制。数据驱动预测控制(DeePC)通过直接利用输入输出轨迹实现无模型控制,但其滚动时域框架需在每个采样时刻求解约束优化问题,对资源受限平台实时部署构成挑战。为此,本文提出一种基于自适应强化学习的事件触发式DeePC(RL-ET-DeePC)框架。训练一个无模型强化学习策略,根据当前系统状态决定是否调用DeePC优化器,从而减少不必要的优化调用,同时保持闭环性能。仿真结果表明,相比周期性DeePC,RL-ET-DeePC可将优化频率降低最多66%,而硬件实验在三维缆绳驱动软体机械臂上验证了零样本迁移能力,优化频率减少34%,跟踪精度与周期性方法相当,且性能更稳定,优于固定阈值的事件触发基线。
原文摘要 · Abstract (English)
Soft robots are challenging to control due to their nonlinear and time-varying dynamics. Data-enabled predictive control (DeePC) offers a model-free alternative by directly leveraging measured input-output trajectories to construct a predictive controller. However, its receding-horizon formulation requires solving a constrained optimization problem at every sampling instant, which can be computationally demanding for real-time deployment on resource-limited robotic platforms. To address this limitation, we propose an adaptive reinforcement-learning-based event-triggered DeePC (RL-ET-DeePC) framework for soft robotic control. A model-free RL policy is trained to determine when to invoke the DeePC optimizer based on the current system state representation, thereby reducing unnecessary optimization calls while preserving closed-loop performance. Simulation results show that RL-ET-DeePC reduces optimization frequency by up to 66% compared to periodic DeePC, while maintaining comparable tracking accuracy. Hardware experiments on a three-dimensional cable-driven soft robotic arm demonstrate zero-shot transfer, achieving a 34% reduction in optimization frequency with tracking accuracy comparable to periodic DeePC and more consistent performance than a static threshold-based event-triggered baseline.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。