arXiv:2606.27353cs.RO2026-06

让机器人持续学习适应环境变化,用实时数据自动修正控制策略。

Continual Robot Policy Learning via Variational Neural Dynamics

论文配图:Continual Robot Policy Learning via Variational Neural Dynamics
图 1 · 摘自论文原文
  • 结合物理先验与神经残差,建模隐藏的动态变化
  • 真实飞行中1秒内恢复风扰,速度比传统方法快5倍
  • 适合需要长期部署的无人机、机械臂等复杂系统

现实世界中的机器人很少在固定动力学下运行:风速变化、负载波动、电池衰减、接触点移动和硬件磨损均会影响表现。然而大多数基于学习的控制器仅在训练时优化一次,部署后无法利用实际经验持续改进。本文提出一种持续学习框架,通过真实状态-动作轨迹学习条件感知的动力学模型,将解析物理先验与神经残差相结合以捕捉未建模效应。一个循环编码器从近期交互中推断当前隐藏状态,并以此条件化残差模型与策略。策略通过可微仿真进行优化,使用从潜在模型中采样的多样化动力学。部署时,采样条件由在线推断的真实交互条件替代,使策略通过识别而非重新拟合来恢复周期性扰动。大量仿真与真实实验表明,该框架在多种未观测扰动下提升了策略性能。在真实四旋翼轨迹跟踪任务中,面对变化风速,策略约1秒内恢复,比在线残差拟合快约5倍;对大扰动下的悬停与跟踪误差分别降低65.7%和53.3%,优于现有在线自适应方法。

原文摘要 · Abstract (English)

Robots deployed in the real world rarely operate under a single fixed dynamics model: wind changes, payloads vary, batteries drain, contacts shift, and hardware wears. Yet most learning-based controllers are trained once and deployed as if learning were complete. This prevents the robot from using deployment experience to further improve task performance. In this work, we propose a continual learning framework that uses real-world experience to improve robot policies under hidden and recurring dynamics. Our method learns a condition-aware dynamics model from real state-action trajectories by combining an analytical physics prior with a neural residual for unmodeled effects. A recurrent encoder infers the current hidden condition from recent interaction, and this estimate conditions both the residual model and the policy. Policy learning is performed via differentiable simulation using diverse learned dynamics sampled from the latent model. At deployment, these sampled conditions are replaced by conditions inferred online from recent real interaction, allowing the policy to recover recurring dynamics by recognition rather than residual re-fitting. Through extensive simulation studies and real-world experiments, we demonstrate that the framework improves policy performance under diverse unobserved disturbances. On real quadrotor trajectory tracking under changing wind, the policy recovers from recurring disturbances in roughly 1s, about 5x faster than online residual re-fitting. It also reduces large-disturbance hover and tracking errors by 65.7% and 53.3% over the state-of-the-art online adaptation approaches

持续学习机器人控制动态建模强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。