用在线滤波预测疲劳,让机器人协作更安全高效
Safe reinforcement learning with online filtering for fatigue-predictive human-robot task planning and allocation in production

- 用粒子滤波实时更新疲劳参数,应对个体差异
- 结合约束双延迟双深度Q网络,确保疲劳不超限
- 适合工业5.0中需人机协同的动态生产场景
人机协同制造是工业5.0的核心,强调人体工学以提升工人福祉。本文针对动态人机任务规划与分配(HRTPA)问题,即在保证工人疲劳处于安全范围内的前提下,决定何时执行任务及由谁完成,以最大化效率。疲劳约束与生产动态叠加使问题复杂度显著上升。传统疲劳恢复模型依赖静态预设超参数,但现实中疲劳敏感性受工作条件变化、睡眠不足等因素影响而日异。为此,我们把疲劳参数视为不确定,并基于生产过程中的疲劳进展实现在线估计。提出PF-CD3Q方法:将粒子滤波与约束双延迟双深度Q学习结合,用于实时疲劳预测的人机任务规划。首先构建基于粒子滤波的疲劳追踪器,实时更新疲劳模型参数;再将其嵌入CD3Q,在决策时进行任务级疲劳预测,排除超出疲劳阈值的任务,从而约束动作空间,将问题建模为约束马尔可夫决策过程(CMDP)。
原文摘要 · Abstract (English)
Human-robot collaborative manufacturing, a core aspect of Industry 5.0, emphasizes ergonomics to enhance worker well-being. This paper addresses the dynamic human-robot task planning and allocation (HRTPA) problem, which involves determining when to perform tasks and who should execute them to maximize efficiency while ensuring workers' physical fatigue remains within safe limits. The inclusion of fatigue constraints, combined with production dynamics, significantly increases the complexity of the HRTPA problem. Traditional fatigue-recovery models in HRTPA often rely on static, predefined hyperparameters. However, in practice, human fatigue sensitivity varies daily due to factors such as changed work conditions and insufficient sleep. To better capture this uncertainty, we treat fatigue-related parameters as inaccurate and estimate them online based on observed fatigue progression during production. To address these challenges, we propose PF-CD3Q, a safe reinforcement learning (safe RL) approach that integrates the particle filter with constrained dueling double deep Q-learning for real-time fatigue-predictive HRTPA. Specifically, we first develop PF-based estimators to track human fatigue and update fatigue model parameters in real-time. These estimators are then integrated into CD3Q by making task-level fatigue predictions during decision-making and excluding tasks that exceed fatigue limits, thereby constraining the action space and formulating the problem as a constrained Markov decision process (CMDP).
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。