通过参考增强学习,提升腱驱动连续机器人在复杂轨迹下的控制精度与稳定性。
Reference-Augmented Learning for Precise Tracking Policy of Tendon-Driven Continuum Robots
- 用可微分RNN构建动力学代理模型,作为优化控制策略的梯度桥梁。
- 在三种多尺度扰动下训练,使策略减少50.9%的位置误差,显著优于传统方法。
- 适合需要高精度、强鲁棒性的连续机器人控制场景,如医疗手术或柔性操作。
腱驱动连续机器人(TDCRs)因其高度非线性、路径依赖的动力学特性及非马尔可夫行为,控制难度大。传统基于雅可比的方法常受滞回效应影响产生振荡,而常规学习方法在分布外轨迹上泛化能力差。本文提出一种参考增强的离线学习框架,实现对TDCRs的6自由度精确跟踪控制。通过使用可微分RNN构建的动力学代理模型作为梯度桥梁,我们基于增强的参考分布优化控制策略。该多尺度增强方案引入随机偏置、谐波扰动和随机游走,迫使策略内化多样化的跟踪误差恢复机制,无需额外硬件交互。在三段式TDCR平台上的实验表明,所提策略相比非增强基线平均位置误差降低50.9%,且在不同速度下均显著优于基于雅可比的方法,在精度和稳定性上表现更优。
原文摘要 · Abstract (English)
Tendon-Driven Continuum Robots (TDCRs) pose significant control challenges due to their highly nonlinear, path-dependent dynamics and non-Markovian characteristics. Traditional Jacobian-based controllers often struggle with hysteresis-induced oscillations, while conventional learning-based approaches suffer from poor generalization to out-of-distribution trajectories. This paper proposes a reference-augmented offline learning framework for precise 6-DOF tracking control of TDCRs. By leveraging a differentiable RNN-based dynamics surrogate as a gradient bridge, we optimize a control policy through an augmented reference distribution. This multi-scale augmentation scheme incorporates stochastic bias, harmonic perturbations, and random walks, forcing the policy to internalize diverse tracking error recovery mechanisms without additional hardware interaction. Experimental results on a three-section TDCR platform demonstrate that the proposed policy achieves a 50.9\% reduction in average position error compared to non-augmented baselines and significantly outperforms Jacobian-based methods in both precision and stability across various speeds. For implementation details and source code, please refer to https://github.com/ZiqingZou/ContinuumControl.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。