提出人机协同自适应模型,确保康复机器人系统稳定收敛到最优交互状态。
Human Machine Co-Adaptation Model and Its Convergence Analysis
- 基于协作马尔可夫决策过程建模人机互动学习
- 证明模型在特定条件下唯一收敛至纳什均衡点
- 设计算法提升多解情形下找到全局最优解的概率
机器人辅助康复的核心在于人机界面设计,需兼顾患者与机器需求。现有设计多聚焦于机器控制算法,常要求患者长时间适应。本文提出基于合作自适应马尔可夫决策过程(CAMDPs)的新方法,揭示交互学习的本质,提供理论支持与实践指导。建立了CAMDPs收敛的充分条件,确保纳什均衡点的唯一性;在此基础上,保障系统收敛至唯一纳什均衡。针对存在多个纳什均衡的情形,提出调整价值评估与策略优化算法的策略,提高收敛至全局最小纳什均衡的概率。数值实验验证了所提条件与算法的有效性,展示了其在实际场景中的适用性与鲁棒性。收敛条件与唯一最优纳什均衡的识别,有助于构建更高效的适应性人机系统,服务于机器人辅助康复。
原文摘要 · Abstract (English)
The key to robot-assisted rehabilitation lies in the design of the human-machine interface, which must accommodate the needs of both patients and machines. Current interface designs primarily focus on machine control algorithms, often requiring patients to spend considerable time adapting. In this paper, we introduce a novel approach based on the Cooperative Adaptive Markov Decision Process (CAMDPs) model to address the fundamental aspects of the interactive learning process, offering theoretical insights and practical guidance. We establish sufficient conditions for the convergence of CAMDPs and ensure the uniqueness of Nash equilibrium points. Leveraging these conditions, we guarantee the system's convergence to a unique Nash equilibrium point. Furthermore, we explore scenarios with multiple Nash equilibrium points, devising strategies to adjust both Value Evaluation and Policy Improvement algorithms to enhance the likelihood of converging to the global minimal Nash equilibrium point. Through numerical experiments, we illustrate the effectiveness of the proposed conditions and algorithms, demonstrating their applicability and robustness in practical settings. The proposed conditions for convergence and the identification of a unique optimal Nash equilibrium contribute to the development of more effective adaptive systems for human users in robot-assisted rehabilitation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。