用强化学习研究反馈在技能习得中的作用,发现学时需多维反馈,但执行时可无反馈。
Using reinforcement learning to probe the role of feedback in skill acquisition
- 用强化学习代理在真实水槽中控制旋转圆柱,通过反馈优化阻力性能。
- 仅数分钟真实交互即学会高性能策略,且无反馈时仍能维持相近表现。
- 学习需丰富反馈,执行可无需反馈,目标不同导致学习条件差异显著。
许多高绩效人类活动几乎无需外部反馈:如花样滑冰运动员完成三周跳、投手投出制胜曲线球、咖啡师制作拉花艺术。为在受控条件下研究技能习得过程,我们绕过人类被试,直接将通用强化学习代理接入桌面循环水道中的旋转圆柱,以最大化或最小化阻力。该系统具备多重优势:一是物理系统,流动高度混沌,难以精确建模或仿真;二是目标明确(阻力增减),奖励函数可直接定义,但优秀策略不明显;三是已有数十年实验研究提供简单高效的开环策略;四是成本低,远比人类实验易复现。实验发现,高维流场反馈使代理仅需几分钟真实交互即可发现高性能阻力控制策略;当后续回放相同动作序列而无反馈时,性能几乎不变。这表明执行已学策略无需反馈,尤其无需流场反馈。令人意外的是,训练中若无流场反馈,代理在阻力最大化任务中完全无法发现有效策略,但在阻力最小化任务中仍能成功,尽管更慢且更不可靠。研究揭示:学习高性能技能所需信息量可能大于执行所需,学习条件的优劣仅取决于目标,而非动力学复杂性或策略复杂性。
原文摘要 · Abstract (English)
Many high-performance human activities are executed with little or no external feedback: think of a figure skater landing a triple jump, a pitcher throwing a curveball for a strike, or a barista pouring latte art. To study the process of skill acquisition under fully controlled conditions, we bypass human subjects. Instead, we directly interface a generalist reinforcement learning agent with a spinning cylinder in a tabletop circulating water channel to maximize or minimize drag. This setup has several desirable properties. First, it is a physical system, with the rich interactions and complex dynamics that only the physical world has: the flow is highly chaotic and extremely difficult, if not impossible, to model or simulate accurately. Second, the objective -- drag minimization or maximization -- is easy to state and can be captured directly in the reward, yet good strategies are not obvious beforehand. Third, decades-old experimental studies provide recipes for simple, high-performance open-loop policies. Finally, the setup is inexpensive and far easier to reproduce than human studies. In our experiments we find that high-dimensional flow feedback lets the agent discover high-performance drag-control strategies with only minutes of real-world interaction. When we later replay the same action sequences without any feedback, we obtain almost identical performance. This shows that feedback, and in particular flow feedback, is not needed to execute the learned policy. Surprisingly, without flow feedback during training the agent fails to discover any well-performing policy in drag maximization, but still succeeds in drag minimization, albeit more slowly and less reliably. Our studies show that learning a high-performance skill can require richer information than executing it, and learning conditions can be kind or wicked depending solely on the goal, not on dynamics or policy complexity.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。