用强化学习选4个最佳姿势,让三自由度踝关节康复机器人校准更高效精准。
D-Optimality-Guided Reinforcement Learning for Efficient Open-Loop Calibration of a 3-DOF Ankle Rehabilitation Robot
- 基于D-最优准则,用强化学习从50个姿态中选出4个最有效的校准姿势。
- 实测表明,新方法使信息矩阵行列式均值提升超两个数量级,方差更小。
- 仅用4个优化姿态的参数估计,比随机选50个更稳定可靠,适合临床部署。
精确对齐多自由度康复机器人对安全有效的患者训练至关重要。本文提出一种两阶段校准框架,针对自研三自由度(3-DOF)踝关节康复机器人。首先,采用基于克罗内克积的开环校准方法,将输入输出对齐转化为线性参数辨识问题,并通过所得信息矩阵定义实验设计目标。在此基础上,将校准姿态选择建模为受D-最优准则指导的组合实验设计问题,即在有限候选集中选取使信息矩阵行列式最大的子集。为实现约束下的实际选择,使用近端策略优化(PPO)代理在仿真中训练,从50个候选姿态中选择4个最具信息量的姿态。仿真与真实机器人测试结果一致显示,所学策略生成的姿态组合显著优于随机选择:PPO实现的信息矩阵行列式均值超过两个数量级,且方差更低。此外,真实世界结果显示,仅用4个D-最优引导姿态识别出的参数向量,其跨实验预测一致性优于由50个无结构姿态获得的估计。该框架在保持鲁棒参数估计的同时提升了校准效率,为多自由度康复机器人的高精度对齐提供了实用指导。
原文摘要 · Abstract (English)
Accurate alignment of multi-degree-of-freedom rehabilitation robots is essential for safe and effective patient training. This paper proposes a two-stage calibration framework for a self-designed three-degree-of-freedom (3-DOF) ankle rehabilitation robot. First, a Kronecker-product-based open-loop calibration method is developed to cast the input-output alignment into a linear parameter identification problem, which in turn defines the associated experimental design objective through the resulting information matrix. Building on this formulation, calibration posture selection is posed as a combinatorial design-of-experiments problem guided by a D-optimality criterion, i.e., selecting a small subset of postures that maximises the determinant of the information matrix. To enable practical selection under constraints, a Proximal Policy Optimization (PPO) agent is trained in simulation to choose 4 informative postures from a candidate set of 50. Across simulation and real-robot evaluations, the learned policy consistently yields substantially more informative posture combinations than random selection: the mean determinant of the information matrix achieved by PPO is reported to be more than two orders of magnitude higher with reduced variance. In addition, real-world results indicate that a parameter vector identified from only four D-optimality-guided postures provides stronger cross-episode prediction consistency than estimates obtained from a larger but unstructured set of 50 postures. The proposed framework therefore improves calibration efficiency while maintaining robust parameter estimation, offering practical guidance for high-precision alignment of multi-DOF rehabilitation robots.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。