用婴儿发育研究方法评估机器人走路,发现训练方式影响行为复杂性。
Evaluating Robots Like Human Infants: A Case Study of Learned Bipedal Locomotion
- 模仿婴儿研究设计强化学习训练方案
- 发现不同训练方式导致行走效率与稳定性差异
- 适合关注机器人行为发育的研究者
通常,学习型机器人控制器通过非系统化训练并以平均累积奖励等粗粒度指标评估。这种方法虽可用于比较学习算法,但难以揭示不同训练方案的影响,也缺乏对学习行为丰富性的理解。相比之下,人类婴儿和动物虽也经历非系统化训练,但发展心理学家会通过高度控制的实验,使用精细指标如行走成功率、速度和前瞻调整来评估其表现。然而,婴儿研究受限于实际操作。本文以仿真双足机器人Cassie为案例,借鉴婴儿学步研究方法,系统设计强化学习训练方案,并在模拟环境中测试,类似婴儿实验但无实际限制。结果揭示了不同训练方案对行为发展的显著影响,且其学习过程与婴儿学步有相似性。该跨学科方法为未来系统研究训练对复杂学习行为发展的影响提供了新思路。
原文摘要 · Abstract (English)
Typically, learned robot controllers are trained via relatively unsystematic regimens and evaluated with coarse-grained outcome measures such as average cumulative reward. The typical approach is useful to compare learning algorithms but provides limited insight into the effects of different training regimens and little understanding about the richness and complexity of learned behaviors. Likewise, human infants and other animals are "trained" via unsystematic regimens, but in contrast, developmental psychologists evaluate their performance in highly-controlled experiments with fine-grained measures such as success, speed of walking, and prospective adjustments. However, the study of learned behavior in human infants is limited by the practical constraints of training and testing babies. Here, we present a case study that applies methods from developmental psychology to study the learned behavior of the simulated bipedal robot Cassie. Following research on infant walking, we systematically designed reinforcement learning training regimens and tested the resulting controllers in simulated environments analogous to those used for babies--but without the practical constraints. Results reveal new insights into the behavioral impact of different training regimens and the development of Cassie's learned behaviors relative to infants who are learning to walk. This interdisciplinary baby-robot approach provides inspiration for future research designed to systematically test effects of training on the development of complex learned robot behaviors.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。